701 to 725 of 788 Permanent Grafana Jobs

SRE-NOC Engineer: Master Incident Response & Reliability

Location
United Kingdom
responsibilities with reliability engineering, focusing on 24/7 service reliability, incident response, and automation. You’ll own runbooks, design alerting, build dashboards with Grafana, and work with cross‐functional teams to reduce toil. The role suits those who engineer solutions rather than only respond to alerts, with a focus ...

Remote UKI Regional Enterprise Growth Director

Location
United Kingdom
Grafana Labs is seeking a Regional Sales Director, Enterprise Growth to lead a team of Enterprise Growth Account Executives across the UK & Ireland. Drive revenue growth, attract and retain talent, and expand the customer base in the region. You will mentor the team, shape strategy, and partner with multiple groups ...

Senior DBA: Lead High-Perf MySQL/PostgreSQL & QuestDB

Location
Greater London, England, United Kingdom
successful candidate will work closely with Engineering and Operations to ensure high-performing, well-managed databases that support our exchange technology products and Grafana dashboards. #J-18808-Ljbffr ...

Security Platform Engineer, UK Security Operations

Hiring Organisation
Google
Location
London, UK
Employment Type
Full-time
Experience with Kubernetes security, including workload isolation, Role-Based Access Control (RBAC), and network policies, containerisation, orchestration, and Kubernetes observability tools (e.g., Falco, Prometheus, Grafana).Experience with infrastructure-as-code and configuration management tools (e.g., Terraform, Helm, ArgoCD).Active, or the ability to obtain, a Developed Vetting (DV) UK security … Experience with Kubernetes security, including workload isolation, Role-Based Access Control (RBAC), and network policies, containerisation, orchestration, and Kubernetes observability tools (e.g., Falco, Prometheus, Grafana).Experience with infrastructure-as-code and configuration management tools (e.g., Terraform, Helm, ArgoCD).Active, or the ability to obtain, a Developed Vetting (DV) UK security ...

Kubernetes Platform Engineer

Location
Greater London, England, United Kingdom
FluxCD) for safe, auditable changes. Drive Infrastructure as Code practices with Terraform and Helm for reliable and repeatable builds. Heavily embed observability using Prometheus, Grafana, and OpenTelemetry to make systems measurable and reliable. Stay ahead of Kubernetes evolution by testing and adopting new versions and features early. Collaborate with teams … Open Policy Agent. Proven ability to troubleshoot complex performance and reliability issues across infrastructure and workloads. Experience with observability tools such as Prometheus, Grafana, and OpenTelemetry to monitor cluster metrics and health. Great communication skills, with experience collaborating with internal platform users to gather feedback and deliver improvements. Experience writing ...

Kubernetes Platform Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
FluxCD) for safe, auditable changes. Drive Infrastructure as Code practices with Terraform and Helm for reliable and repeatable builds. Heavily embed observability using Prometheus, Grafana, and OpenTelemetry to make systems measurable and reliable. Stay ahead of Kubernetes evolution by testing and adopting new versions and features early. Collaborate with teams … Open Policy Agent. Proven ability to troubleshoot complex performance and reliability issues across infrastructure and workloads. Experience with observability tools such as Prometheus, Grafana and OpenTelemetry to monitor cluster metrics and health. Great communication skills, with experience collaborating with internal platform users to gather feedback and deliver improvements. Experience writing ...

Network Automation & OSS Designer

Location
Greater London, England, United Kingdom
Architect AIOps capabilities including closed‐loop automation, anomaly detection, and predictive analytics for network operations. Integrate OSS observability with cloud‐native monitoring stacks (Prometheus, Grafana, OpenTelemetry, Elasticsearch). Lead design of intent‐based networking and policy‐driven automation frameworks. Collaborate with product managers, network engineers, platform teams, and DevOps practitioners …/CNF lifecycle management). Understanding of AIOps platforms and closed‐loop automation design for network operations. Experience with observability tooling: OpenTelemetry, Prometheus, Grafana, Jaeger, Loki, ELK Stack. Knowledge of ML/AI model integration for anomaly detection, root‐cause analysis, and predictive network management. Strong grasp of TM Forum ...

ML/AI Engineer

Location
Manchester, England, United Kingdom
observability for models and pipelines: drift, data quality, fairness signals, latency, GPU utilisation, error budgets, and SLOs/SLIs via Prometheus, Grafana, and Dynatrace. Establish actionable alerting and runbooks for on‐call operations; drive incident reviews and reliability improvements. Operate a model registry (e.g., MLflow) with experiment tracking, versioning, lineage … pipelines; experience with GitOps, artefact repositories, and environment promotion. Practical experience with CUDA, TensorRT, Triton, TorchServe, and GPU scheduling/optimisation. Proficiency in Prometheus, Grafana, Dynatrace defining SLIs/SLOs and alert thresholds for ML systems. Experience operating MLflow (or equivalent) for experiment tracking, model bundling, and deployments. Expert ...

Site Reliability Engineer with Python

Hiring Organisation
BC Forward
Location
Charlotte, North Carolina, United States
Employment Type
Permanent
Salary
USD Hourly
Job Title: Site Reliability Engineer with Python Location: Pennington, NJ Duration: Contract - 9 months Pay Range: $73.67/hr (W2) Job ID: 410592 About BCforward BCforward is a leading global IT consulting and workforce solutions ...

Senior Infrastructure Engineer

Location
Greater London, England, United Kingdom
Thought Machine's mission is bold – to properly and permanently rid the world's banks of legacy technology. To achieve this, we have developed the foundations of modern banking through core and payments technology which ...

Site Reliability Engineer

Hiring Organisation
Proactive Appointments
Location
Gloucester, Gloucestershire, UK
Employment Type
Full-time
Job Description Site Reliability Engineer - DV Cleared Our client is urgently looking for an experienced Site Reliability Engineer to join their team on a contract basis, initially for 6 months with a view to extend. ...

Software Engineer, Security Rules

Location
Greater London, England, United Kingdom
About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world's largest networks that powers millions of websites and other Internet properties ...

Site Reliability Engineer – 11863CF

Location
Greater London, England, United Kingdom
11863CF £500 – 580 per day Site Reliability Engineer – DV Cleared Our client is urgently looking for an experienced Site Reliability Engineer to join their team on a contract basis, initially for 6 months with a ...

Scala Developer

Location
Sunderland, England, United Kingdom
Senior Scala/JVM Software Developer Newcastle, Tyne & Wear, United Kingdom - £400 - £450 Inside IR 35 per day Contract About Scrumconnect Consulting Scrumconnect Consulting is a multi-award-winning digital consultancy, recognised for delivering impactful ...

Staff Product Manager, OpenTelemetry | UK | Remote

Location
United Kingdom
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. We’re scaling fast and staying true to what makes us different: an open-source legacy, a global collaborative culture, and a passion ...

Client Implementation Engineer

Hiring Organisation
Bright Purple Resourcing
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£45,000
natural problem solver with a curious mind, capable of grappling with difficult technical challenges. You'll be working with Linux, Bash, Python and Grafana, onboarding customers, integrating solutions into their environments and investigating production issues alongside an experienced engineering team. There's plenty of variety, with opportunities to influence … from legacy Java components Work directly with clients on technical onboarding, implementation, issue resolution and product training Build and support dashboards and visualisations in Grafana, and explore operational data in systems like QuestDB and InfluxDB Collaborate with software engineers and internal ML teams to shape new product capabilities and deploy ...

Senior Site Reliability Engineer - Infrastructure and Agentic Automation

Hiring Organisation
CYNET SYSTEMS
Location
Santa Clara, California, United States
Employment Type
Permanent
Salary
USD Hourly
such as Claude Code, Codex CLI, and Model Context Protocol) to automate diagnostics, ticket triage, and operational remediation workflows. Implement comprehensive observability platforms (Datadog, Grafana, custom data pipelines) to monitor fleet health, track Chef/Cinc run metrics, and proactively surface system anomalies. Partner with Windows and Linux engineering teams … workflows leveraging AI agent tooling, LLM orchestration, or agentic frameworks. Expertise in configuring end-to-end monitoring, metrics collection, logging, and alerting (Datadog/Grafana) to ensure platform reliability. Experience managing and securing Windows infrastructure alongside Linux. Proficiency in languages such as Python, Go, PowerShell, or Bash for automation ...

Platform Site Reliability Engineer

Location
Gloucester, England, United Kingdom
tooling for our support organisation Apply ITSM frameworks: Incident, Major Incident, Change Management, and service improvement. Maintain and enhance Radiant’s observability stack: Prometheus, Grafana, and custom monitoring integrations Operate and support services in 24x7 production environments, including on-call rotation Contribute to Incident postmortem analyses, root cause analysis, document … switching Strong experience with API interrogation Strong experience with infrastructure scripting and automation (Bash, Python, Ansible) Deep understanding of observability principles and tools (Prometheus, Grafana preferred) Strong grasp of ITSM and service operation best practices Excellent communication and mentorship skills Comfortable interfacing with internal stakeholders and external customers Bonus: Knowledge ...

Infrastructure Site Reliability Engineer

Location
Gloucester, England, United Kingdom
tooling for our support organisation Apply ITSM frameworks: Incident, Major Incident, Change Management, and service improvement. Maintain and enhance ’s observability stack: Prometheus, Grafana, and custom monitoring integrations Operate and support services in 24x7 production environments, including on-call rotation Contribute to Incident postmortem analyses, root cause analysis, document learnings … DHCP, VLANs, routing, switching Strong experience with infrastructure scripting and automation (Bash, Python, Ansible) Deep understanding of observability principles and tools (Prometheus, Grafana) Hands-on experience operating orchestration platforms (Kubernetes, MAAS, Tinkerbell) Strong grasp of ITSM and service operation best practices Excellent communication and mentorship skills Comfortable interfacing with internal ...

Platform Engineer

Hiring Organisation
Springer Nature
Location
London, UK
Employment Type
Full-time
Platform, as well as managing an internal database platform that runs over 1,200 database servers. Out tech stack focuses on: Kubernetes, Kubevela, GKE Grafana LGTM stack, Sentry, Open Telemetry Concourse, GitHub Actions Terraform, Ansible Postgres, Mongo, MySQL GCP, AWSRole Responsibilities: Meet your delivery and participation responsibilities as a contributor … language features and idiomatic practices. You can demonstrate your experience with some of the core technologies in use (Kubernetes, Kubevela, GKE, EKS, Grafana LGTM stack, Sentry, Open Telemetry, Concourse, GitHub Actions, Terraform, Ansible, Postgres, Mongo) You have experience in building APIs, exposing components in a programmable way, to be consumed ...

Senior AI Quality Engineer

Location
Greater London, England, United Kingdom
back on ambiguous acceptance criteria, and surface risk before code is written Close the loop on production issues using our observability stack (Datadog, Sentry, Grafana) - tying test coverage back to real customer impact Ensure teams have Service Level Objectives set up and are achieving them Run targeted exploratory testing … Kotlin, AWS, Postgres, RabbitMQ, Docker, Kubernetes Testing: PHPUnit, Behat, JUnit, Kotest, Jest, Maestro, K6 Tooling and observability: GitHub, GitHub Actions, Jira, Confluence, Datadog, Sentry, Grafana Why join? See your work matter: Our products are used by millions of customers - the quality bar you help set has a direct line ...

Sr. FinOps Engineer (68022) (DEAI DS) Cloud & Data Engineering United Kingdom

Location
Greater London, England, United Kingdom
maintain FinOps dashboards and reports: executive cost views, business-unit breakdowns, trend analysis, anomaly alerts, and licence utilisation reports using Power BI, Tableau, or Grafana Conduct usage, pattern, and licence utilisation analysis: identify underused licences, oversized hardware, idle resources, peak/off-peak patterns, and capacity rebalancing opportunities Support … skills: SQL (advanced), Python or PowerShell for data extraction, transformation, and automation; experience with ETL/ELT pipelines Dashboard development: Power BI (preferred), Tableau, Grafana, or equivalent BI tools — building interactive dashboards with drill-down, filtering, and alerting Experience with on-premises infrastructure: server/VM inventory, storage systems, network ...

Sr. FinOps Engineer (68022)

Location
Greater London, England, United Kingdom
maintain FinOps dashboards and reports: executive cost views, business‐unit breakdowns, trend analysis, anomaly alerts, and licence utilisation reports using Power BI, Tableau, or Grafana Conduct usage, pattern, and licence utilisation analysis: identify underused licences, oversized hardware, idle resources, peak/off‐peak patterns, and capacity rebalancing opportunities Support … skills: SQL (advanced), Python or PowerShell for data extraction, transformation, and automation; experience with ETL/ELT pipelines Dashboard development: Power BI (preferred), Tableau, Grafana, or equivalent BI tools — building interactive dashboards with drill‐down, filtering, and alerting Experience with on‐premises infrastructure: server/VM inventory, storage systems, network ...

Software Engineer, Security Rules

Location
Greater London, England, United Kingdom
Cloudflare ·United States, Lisbon, Portugal/London, United Kingdom Job Description About Us At Cloudflare, we are on a mission to help build a better Internet. Today the company runs one of the world’s ...

SRE Engineer

Location
Greater London, England, United Kingdom
Our client is looking for a SRE Engineer combining software and IT engineering principles to build and maintain reliable, scalable, and high-performing systems. Job Responsibilities: Scope technical projects and break them down into user ...