251 to 275 of 297 Remote/Hybrid Grafana Jobs

Platform Engineer

Location
Greater London, England, United Kingdom
huge, distributed scale efficiently Monitoring and alerting: Measuring application performance and delivering insights, metrics and relevant alerts to the engineering teams with ELK, Grafana and New Relic Ownership: Driving engineering teams to own their infrastructure and costs by building great tooling, visibility and documentation Security: Setting the standards for fine … understanding of CI/CD and relevant tooling (we use GitHub Actions and Argo CD) Expertise in logging and monitoring at scale (e.g.S3, Graphite, Grafana, ELK, NewRelic, Datadog) Knowledge of a DevOps toolchain to drive ownership of a self-hosted platform Competent in Git and the GitOps philosophy Familiarity with ...

Senior Infrastructure Engineer (GCP) - Engine by Starling

Location
Cardiff, Wales, United Kingdom
workloads and CI/CD Experience with observability tooling — Cloud Monitoring, Cloud Logging, Cloud Trace, Managed Service for Prometheus and OpenTelemetry (we also use Grafana) Experience setting up Google Workspace/Google Cloud Identity Experience with automation using a scripting language like Python or Go Experience implementing CI/… native Container-based architecture Kubernetes (GKE on GCP, EKS on AWS) TeamCity for CI/CD (with multiple production releases per day) Terraform and Grafana RDS and CloudSQL for PostgreSQL Our Interview Process Interviewing is a two-way process and we want you to have the time and opportunity ...

Senior Infrastructure Engineer (GCP) - Engine by Starling

Location
Greater London, England, United Kingdom
workloads and CI/CD Experience with observability tooling — Cloud Monitoring, Cloud Logging, Cloud Trace, Managed Service for Prometheus and OpenTelemetry (we also use Grafana) Experience setting up Google Workspace/Google Cloud Identity Experience with automation using a scripting language like Python or Go Experience implementing CI/… native Container-based architecture Kubernetes (GKE on GCP, EKS on AWS) TeamCity for CI/CD (with multiple production releases per day) Terraform and Grafana RDS and CloudSQL for PostgreSQL Our Interview Process Interviewing is a two-way process and we want you to have the time and opportunity ...

Senior Java Full Stack Engineer

Hiring Organisation
Pinnacle Technical Resources
Location
Tempe, Arizona, United States
Employment Type
Permanent
Salary
USD 55 Hourly
Position: Sr. Java Full Stack Developer Location: Tempe, Arizona Duration: Contract Job ID: 180092 Interview: F2F required Job Overview: We are seeking a highly skilled and experienced Sr. Java Full Stack Developer to join our ...

Engineering Manager, App Security (Cloud & OSS) - Remote

Location
United Kingdom
Grafana Labs is seeking an Engineering Manager for Application Security in the EMEA region (Remote, UK). You will lead security engineering squads, drive automation, and shape security posture across cloud and on‐premise components. The role emphasizes coaching, risk‐based prioritization, and collaboration with product and engineering teams. ...

Remote UKI Regional Enterprise Growth Director

Location
United Kingdom
Grafana Labs is seeking a Regional Sales Director, Enterprise Growth to lead a team of Enterprise Growth Account Executives across the UK & Ireland. Drive revenue growth, attract and retain talent, and expand the customer base in the region. You will mentor the team, shape strategy, and partner with multiple groups ...

ML/AI Engineer

Location
Manchester, England, United Kingdom
observability for models and pipelines: drift, data quality, fairness signals, latency, GPU utilisation, error budgets, and SLOs/SLIs via Prometheus, Grafana, and Dynatrace. Establish actionable alerting and runbooks for on‐call operations; drive incident reviews and reliability improvements. Operate a model registry (e.g., MLflow) with experiment tracking, versioning, lineage … pipelines; experience with GitOps, artefact repositories, and environment promotion. Practical experience with CUDA, TensorRT, Triton, TorchServe, and GPU scheduling/optimisation. Proficiency in Prometheus, Grafana, Dynatrace defining SLIs/SLOs and alert thresholds for ML systems. Experience operating MLflow (or equivalent) for experiment tracking, model bundling, and deployments. Expert ...

Scala Developer

Location
Sunderland, England, United Kingdom
Senior Scala/JVM Software Developer Newcastle, Tyne & Wear, United Kingdom - £400 - £450 Inside IR 35 per day Contract About Scrumconnect Consulting Scrumconnect Consulting is a multi-award-winning digital consultancy, recognised for delivering impactful ...

Staff Product Manager, OpenTelemetry | UK | Remote

Location
United Kingdom
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. We’re scaling fast and staying true to what makes us different: an open-source legacy, a global collaborative culture, and a passion ...

Client Implementation Engineer

Hiring Organisation
Bright Purple Resourcing
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£45,000
natural problem solver with a curious mind, capable of grappling with difficult technical challenges. You'll be working with Linux, Bash, Python and Grafana, onboarding customers, integrating solutions into their environments and investigating production issues alongside an experienced engineering team. There's plenty of variety, with opportunities to influence … from legacy Java components Work directly with clients on technical onboarding, implementation, issue resolution and product training Build and support dashboards and visualisations in Grafana, and explore operational data in systems like QuestDB and InfluxDB Collaborate with software engineers and internal ML teams to shape new product capabilities and deploy ...

Senior AI Quality Engineer

Location
Greater London, England, United Kingdom
back on ambiguous acceptance criteria, and surface risk before code is written Close the loop on production issues using our observability stack (Datadog, Sentry, Grafana) - tying test coverage back to real customer impact Ensure teams have Service Level Objectives set up and are achieving them Run targeted exploratory testing … Kotlin, AWS, Postgres, RabbitMQ, Docker, Kubernetes Testing: PHPUnit, Behat, JUnit, Kotest, Jest, Maestro, K6 Tooling and observability: GitHub, GitHub Actions, Jira, Confluence, Datadog, Sentry, Grafana Why join? See your work matter: Our products are used by millions of customers - the quality bar you help set has a direct line ...

Senior Network Engineer

Location
Greater London, England, United Kingdom
Using BGP, EVPN-VXLAN, JunOS (Juniper QFX/MX), OSPF, Spine-Leaf/IP Fabric, VLANs, VRFs, Linux networking, Ansible, Terraform, Prometheus/Grafana, Observability tooling The adventures that await you after becoming Senior Network Engineer at Hack The Box: Design and implement spine-leaf network architectures across multiple data … JunOS Develop and maintain network automation using Ansible, Terraform or similar infrastructure‐as‐code tooling Establish and improve network monitoring, alerting and observability (Prometheus, Grafana, SNMP, streaming telemetry) Plan and execute network capacity upgrades, site bring‐ups and hardware refresh cycles Collaborate with the platform/systems engineering team ...

Senior Network Engineer, Studios

Location
Greater London, England, United Kingdom
through CI pipelines (Python, Ansible, Git); eliminate hand‐edits. Observability: develop monitoring built on streaming telemetry (gNMI/gRPC), flow analysis and modern tooling (Grafana, Zabbix class), serving the 24/7 operations teams as your internal customers. Modernisation: lead the shift from manual workflows to automated NetOps across … history. IPAM/DCIM: experience with NetBox or a similar platform for infrastructure documentation and management; NetBox preferred. Observability: practical experience with modern monitoring (Grafana, Zabbix, ELK class), flow analysis and streaming telemetry. Security: proven experience operating and securing enterprise firewall estates (Palo Alto, Fortinet or Cisco class) and managing ...

Scala Data Engineer - Contract - SC Clearance - London

Hiring Organisation
CBSbutler Holdings Limited trading as CBSbutler
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£530 - £580/day insideIR35
Senior Scala Data Engineer - 6 Month Contract Location: London - Hybrid Duration: 6 months Rate: £530 - £580 per day insideIR35 Clearance: SC Clearance required We are seeking an experienced Senior Scala Data Engineer with active SC ...

Senior Site Reliability Engineer (Observability)

Location
United Kingdom
engineering team better at running its own services. Today our telemetry lives in Azure Monitor, Log Analytics and Application Insights, with Azure Managed Grafana for dashboards and alerting. It works, but it has grown organically. Alert thresholds are inherited rather than designed, retention and cost are not governed, instrumentation … target architecture and standards for the platform - what we instrument, how, where it lands, how long we keep it and what it costs. Make Grafana the place engineers go to understand production: dashboards and alerting designed around services and customer journeys, not around whichever metrics happened to be available. Establish ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. Proficiency in shell scripting for automation and system management. Experience with Infrastructure as Code, including Terraform and Ansible. Knowledge of Cloudflare … operational consistency. Write and contribute to code, telemetry and instrumentation that improve service reliability and observability. Build dashboards and operational views using telemetry from Grafana, Splunk, New Relic and related platforms. Configure and manage Cloudflare edge services using Infrastructure as Code and integrate edge telemetry with observability platforms. Diagnose incidents ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Stoke-On-Trent, Staffordshire, West Midlands, United Kingdom
Employment Type
Permanent, Work From Home
understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. Proficiency in shell scripting for automation and system management. Experience with Infrastructure as Code, including Terraform and Ansible. Knowledge of Cloudflare … operational consistency. Write and contribute to code, telemetry and instrumentation that improve service reliability and observability. Build dashboards and operational views using telemetry from Grafana, Splunk, New Relic and related platforms. Configure and manage Cloudflare edge services using Infrastructure as Code and integrate edge telemetry with observability platforms. Diagnose incidents ...

Senior Backend Engineer - Asset Sales

Location
Greater London, England, United Kingdom
C# stack : Distributed C# and .NET microservices Cloud & orchestration : Hosted on Azure using Kubernetes Architecture : Event-driven, supporting products used at significant scale Observability : Grafana, Azure Application Insights, logs, traces, and metrics AI tooling : Claude and other AI tools used throughout the engineering workflow — design exploration, code generation and review … have experience with similar messaging technology (Kafka, RabbitMQ, etc.) Microservices architecture experience, ideally on Azure and Kubernetes Experience using observability tooling (e.g. Grafana, Application Insights) to understand production behaviour Comfort using AI tools as part of a daily engineering workflow (design, code review, testing, incident investigation) DevOps culture mindset, including ...

Database Reliability Engineer

Location
Manchester, England, United Kingdom
Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will leverage Java-based data collection pipelines and metrics platforms (Prometheus, Grafana, Dash0, Sentry, Humio) to detect performance regressions, slow queries, and health issues before they impact customers Fortify Business Continuity (BCP): Design and implement rigorous Business … region data platforms—while ensuring data integrity, clean relational modeling, and mobility A Security & Observability Mindset: You focus on building deep observability (Prometheus/Grafana/Dash0/Sentry/Humio) and automated security guardrails directly into your Java services so the fleet is secure and monitored by design Interview ...

Database Reliability Engineer

Location
Cardiff, Wales, United Kingdom
Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will leverage Java-based data collection pipelines and metrics platforms (Prometheus, Grafana, Dash0, Sentry, Humio) to detect performance regressions, slow queries, and health issues before they impact customers Fortify Business Continuity (BCP): Design and implement rigorous Business … region data platforms—while ensuring data integrity, clean relational modeling, and mobility A Security & Observability Mindset: You focus on building deep observability (Prometheus/Grafana/Dash0/Sentry/Humio) and automated security guardrails directly into your Java services so the fleet is secure and monitored by design Interview ...

Database Reliability Engineer

Location
Southampton, England, United Kingdom
Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will leverage Java-based data collection pipelines and metrics platforms (Prometheus, Grafana, Dash0, Sentry, Humio) to detect performance regressions, slow queries, and health issues before they impact customers Fortify Business Continuity (BCP): Design and implement rigorous Business … region data platforms—while ensuring data integrity, clean relational modeling, and mobility A Security & Observability Mindset: You focus on building deep observability (Prometheus/Grafana/Dash0/Sentry/Humio) and automated security guardrails directly into your Java services so the fleet is secure and monitored by design Interview ...

Database Reliability Engineer

Location
Greater London, England, United Kingdom
Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will leverage Java-based data collection pipelines and metrics platforms (Prometheus, Grafana, Dash0, Sentry, Humio) to detect performance regressions, slow queries, and health issues before they impact customers Fortify Business Continuity (BCP): Design and implement rigorous Business … region data platforms—while ensuring data integrity, clean relational modeling, and mobility A Security & Observability Mindset: You focus on building deep observability (Prometheus/Grafana/Dash0/Sentry/Humio) and automated security guardrails directly into your Java services so the fleet is secure and monitored by design Interview ...

Technical Product Manager (Superapp)

Location
Greater London, England, United Kingdom
About LendableLendable is on a mission to build the world's best technology to help people get credit and save money. We're building one of the world’s leading fintech companies and are off ...

Site Reliability Engineer

Location
United Kingdom
understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. Proficiency in shell scripting for automation and system management. Experience with Infrastructure as Code, including Terraform and Ansible. Knowledge of Cloudflare … operational consistency. Write and contribute to code, telemetry and instrumentation that improve service reliability and observability. Build dashboards and operational views using telemetry from Grafana, Splunk, New Relic and related platforms. Configure and manage Cloudflare edge services using Infrastructure as Code and integrate edge telemetry with observability platforms. Diagnose incidents ...

Site Reliability Engineer

Location
Manchester, England, United Kingdom
Service Level Objectives (SLO's) for reliability and customer satisfaction. Knowledge of contemporary observability tools, techniques and best practice including Splunk, New Relic, Grafana and PagerDuty. Proficiency in shell scripting for automation and system management tasks. Experience with Infrastructure as Code (IaC), automation and orchestration tools such as Ansible … observability of services, including telemetry, operational APIs and tooling. Building sophisticated dashboards using a range of telemetry data and dash boarding technologies like Grafana, Splunk and New Relic. Actively participating in live incident resolution and post‐mortem analysis, providing effective remediation strategies to improve overall system health and prevent future ...