676 to 700 of 925 Permanent Grafana Jobs

NOC Engineer / SRE

Location
United Kingdom
health monitoring across infrastructure, applications, and dependencies Design and maintain alerting strategies that align with SLIs/SLOs Build dashboards using tools such as: Grafana Reliability Engineering & Automation Automate repetitive operational tasks to reduce manual toil Improve mean time to detect (MTTD) and mean time to resolve (MTTR) Develop scripts ...

Senior Engineering Manager, Developer Experience

Location
Greater London, England, United Kingdom
track record of owning a technical domain end-to-end. You bring strong technical foundation across the DevEx stack: CI/CD, observability (Prometheus, Grafana, or equivalent), Kubernetes-based platforms - sufficient to make sound architectural decisions and earn engineer trust. You know how to lead through ambiguity and organisational change ...

Engineering Team Lead

Location
Cardiff, Wales, United Kingdom
Typescript). Experience working with modern Continuous Integration tooling, such as Jenkins. Experience debugging and troubleshooting issues using monitoring and observability tooling, such as Grafana, Prometheus or equivalent. Values Teamwork Merit Develop Honest Impactful Integrity Benefits Discretionary bonus 25 days annual leave (plus bank holidays) Opportunity to purchase additional annual ...

Vice President, Production Services Application Support

Hiring Organisation
BNY
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
management, post-incident reviews, root cause analysis, and the automation of operational tasks. Experience of using monitoring and observability tools such as AppDynamics, Splunk, Grafana, Cloudprober, and Moogsoft. Knowledge of multi-tier application architecture. Experience of web technologies and internet-based applications. Experience of working with production and non-production ...

Forward Deployed Engineer

Location
Greater London, England, United Kingdom
decisions, and cultural foundations Our Tech Stack Frontend: React, Typescript Backend: Golang & Rust Cloud: Cloudflare, GCP, AWS, multiple LLM providers Tooling: Github Actions, OTEL, Grafana, Terraform, and more How we hire Fill out a short form and hop on a brief intro call with our recruiting team. Complete a quick ...

HPC Systems & Scheduling Developer (m/f/d)

Hiring Organisation
Gesellschaft für wissenschaftliche Datenverarbeitung mbH Göttingen
Location
Göttingen, Niedersachsen, Germany
Employment Type
Permanent
Salary
EUR Annual
workloads Experience enabling optimal, resource-efficient operation of hybrid classical/quantum workloads Familiarity with monitoring and observability tooling for HPC systems (e.g., Prometheus, Grafana) Background in systems administration or DevOps for research computing infrastructure German B2 or higher Our offer Flexible working hours and opportunity for partial mobile working ...

HPC Network Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, UK
Employment Type
Full-time
Linux administration fundamentals: you can debug from the host side as well as the switch sideExperience with network telemetry and monitoring (e.g. Prometheus/Grafana, sFlow/IPFIX, streaming telemetry)Experience running corporate/campus networks: wired and wireless, switching, NAC/802.1X, VPN and remote accessClear communicator who enjoys ...

HPC Network Engineer

Location
Greater London, England, United Kingdom
administration fundamentals: you can debug from the host side as well as the switch side Experience with network telemetry and monitoring (e.g. Prometheus/Grafana, sFlow/IPFIX, streaming telemetry) Experience running corporate/campus networks: wired and wireless, switching, NAC/802.1X, VPN and remote access Clear communicator ...

Infrastructure Engineer

Location
Greater London, England, United Kingdom
team that supports mission-critical production systems and back-office core infrastructure functions. Performance Management: Experience with hardware/network monitoring tools (e.g., Prometheus, Grafana, SNMP) and performance tuning. Professional Communication: Aptitude to maintain a professional and clear relationship in all collaborative technical interactions, inquiries, and status updates. Organisation: Strong ...

Technical Support Engineer

Location
Greater London, England, United Kingdom
C#, Java Script or shell scripting Familiarity with Web Technologies, HTTP and HTTPS Experience with REST, JSON, SOAP Experience with SQL Familiarity with Splunk, Grafana Experience in e-commerce, payment, fin-tech industry is a plus Equal opportunity Airwallex is proud to be an equal opportunity employer. We value diversity ...

Technical Account Manager (UK, EMEA)

Location
Greater London, England, United Kingdom
Have Experience with SCA, dependency management, or security scanning tools Contributions to open-source projects Tools You’ll Use Slack, GitHub, Zendesk, Vitally, HubSpot, Grafana Our Interview Process Informational call with someone from our Talent team Virtual f2f with our Head of Customer Engineering Virtual f2f with a peer Presentation ...

Software Engineer II - Backend (Ruby)

Location
Greater London, England, United Kingdom
systems (e.g. you know your way around logs, traces, metrics, feature flags, and can debug live traffic issues) Familiarity with observability tooling (e.g. Prometheus, Grafana, Sentry and Lightstep) A good understanding of Domain Driven Design experience integrating with third‐party APIs and services, particularly those with nuanced state transitions (like ...

Incident Manager

Location
Leeds, England, United Kingdom
senior technical monitoring role. Deep hands‐on expertise with enterprise monitoring and observability platforms such as Zabbix, Splunk, DynaTrace, New Relic, Datadog, Grafana, or similar tools. Proven experience designing and implementing monitoring strategies for large-scale, distributed, customer-facing digital platforms. Strong understanding of alert engineering principles including threshold tuning ...

Staff / Senior Staff Security Engineer, Vinted Pay

Location
Greater London, England, United Kingdom
written and spoken English. Advantage: experience building and running systems at massive scale (2+ million requests per minute), deep knowledge of observability tooling (Kibana, Grafana, Prometheus), and a passion for introducing new practices. Advantage: AWS security depth (IAM, KMS, multi‐account and multi‐region architecture), or hands‐on CSPM (e.g. ...

Mid-Level Backend Engineer - Driver Squad

Location
Greater London, England, United Kingdom
personal projects or open source contributions. Technologies We Use Golang AWS, CDK (TypeScript), Lambda, SQS, EventBridge, RDS (PostgreSQL) React, React Native GitHub, GitHub Actions Grafana, Loki Claude Code Event‐driven architecture and domain‐driven design Interview Process Intro call with the hiring manager Live coding challenge solving an optimisation problem ...

Production Support Engineer II - UK

Hiring Organisation
Marqeta
Location
London, United Kingdom
Salary
£ 60 K
Linux environmentIntermediate SQL knowledge (relational database experience preferred)Scriptwriting - Python, Ruby, Shell, etcExperience with logging and monitoring tools such as Kibana, Splunk, AppDynamic, SumLogic, Grafana, Datadog, and New RelicThe ability and desire to learn new technologies and toolsExperience in payments and/or accounting systemsExperience working at a high-growth ...

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
will drive the migration of applications from existing monitoring tools (Geneos ITRS, Prometheus, ELK, Splunk, AppDynamics, etc.) to Google Cloud Observability (GCO) and Grafana using OpenTelemetry (OTel) as the instrumentation standard. You will act as a hands‐on technical authority, authoring reusable deployment solutions, configuring telemetry collectors, and providing direct … improvement and automation. Collaborative Enablement: Partner with development and SRE teams to drive the adoption of OpenTelemetry (OTel) and Google Cloud Observability (GCO) and Grafana standards. Regulatory Compliance: Operate effectively within a highly regulated environment, ensuring all observability and deployment solutions comply with relevant enterprise standards and security requirements. Resiliency ...

Staff Field Engineer | UK | Remote

Hiring Organisation
Grafana Labs
Location
United Kingdom
Salary
£ 60 K
Grafana Labs is the company behind Grafana Cloud, the fully managed observability platform trusted by more than 10,000 organizations to ensure reliability, resolve incidents faster, and optimize telemetry at scale. Built on open source and open standards and designed for interoperability across any stack, Grafana Cloud brings … their disparate data, wherever it lives, and move at the speed of their ambitions. Customers, including Anthropic, Bloomberg, NVIDIA, Microsoft, and Salesforce, rely on Grafana Labs. We are a 100% remote company with team members across 40+ countries, backed by leading investors including Lightspeed Venture Partners, Sequoia Capital, GIC, Coatue ...

Senior DevOps Engineer – Build Pipelines

Location
Greater London, England, United Kingdom
Integrate pipeline and deployment processes with appropriate monitoring and observability tooling. Help teams understand build and deployment health. Work with technologies such as: Prometheus Grafana Datadog Establish useful metrics and visibility around pipeline execution and deployments. Engineering Enablement Work directly with application engineering teams to implement new standards. Provide practical … Microsoft Azure Windows Server Windows VMs IIS On-premise infrastructure Hybrid environments Automation & Scripting PowerShell Python Bash or other scripting languages beneficial Observability Prometheus Grafana Datadog Security & Software Supply Chain SonarQube Snyk or equivalent SAST DAST Dependency/CVE scanning Image/container signing Software supply-chain security Application Technologies ...

Senior Engineer, Platform Access

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 70 K
workload management tools, such as IBM Spectrum LSF or comparable schedulers, and enterprise engineering access platforms.Experience implementing monitoring, logging, and observability solutions using Dynatrace, Grafana, ELK/OpenSearch, Splunk, CloudWatch, or similar tools.Understanding of SRE and modern operations practices, including incident resolution, capacity management, AIOps, policy-based automation, self-healing …/CD, and REST APIs, including Infrastructure-as-Code, configuration management, and automated deployment practices.Experience implementing monitoring, logging, and observability solutions using Dynatrace, Grafana, ELK/OpenSearch, Splunk, CloudWatch, or similar tools.Understanding of production support practices, including incident response, handling blocking issues, root-cause analysis, documentation, and continuous service improvement.Familiarity ...

Senior Engineer, Platform Access

Location
Cambridge, England, United Kingdom
management tools, such as IBM Spectrum LSF or comparable schedulers, and enterprise engineering access platforms. Experience implementing monitoring, logging, and observability solutions using Dynatrace, Grafana, ELK/OpenSearch, Splunk, CloudWatch, or similar tools. Understanding of SRE and modern operations practices, including incident resolution, capacity management, AIOps, policy-based automation, self …/CD, and REST APIs, including Infrastructure-as-Code, configuration management, and automated deployment practices. Experience implementing monitoring, logging, and observability solutions using Dynatrace, Grafana, ELK/OpenSearch, Splunk, CloudWatch, or similar tools. Understanding of production support practices, including incident response, handling blocking issues, root-cause analysis, documentation, and continuous ...

Senior DevOps Engineer (AWS)

Hiring Organisation
Duetto
Location
United Kingdom
Salary
£ 50 K
maintain CI/CD pipelines, deployment automation, and infrastructure-as-code using Terraform, Terragrunt, and Chef.Monitor system health using enterprise observability tools (Prometheus, Grafana, DataDog), defining metrics and response playbooks that get ahead of incidents before they happen.Lead security-first practices across the infrastructure estate — proactively addressing vulnerabilities and driving … Artifactory, and Jenkins, with familiarity with GitOps methodologies.Experience with container and orchestration technologies — Docker, ECS, or EKS.Strong working knowledge of enterprise monitoring platforms — Prometheus, Grafana, or DataDog.Solid troubleshooting skills and a track record of leading root cause analyses and implementing durable fixes.Proficiency reading Java, Ruby, Bash/Zsh, HCL, Python ...

Site Reliability Engineer

Hiring Organisation
GoCardless
Location
London, United Kingdom
Salary
£ 70 K
maintenance, improvements, and support for audits and compliance.Tech stack and tools:Python, Ruby, Golang;Terraform;Atlantis AWS, GCP;Kubernetes, GKE;Github, GitHub Actions, ArgoCD;Grafana, Prometheus;DatadogWhat excites you:We use a wide range of technologies, and will never expect you to have experience with all of them. … considered an advantage;Experience with CI/CD tooling such as Github, GitHub Actions, ArgoCD, etc.;Experience working with monitoring tooling such as Grafana, Prometheus, etc.;Experience with relational databases and other datastores, especially around high availability and performance optimisation;Awareness of DevOps and Agile principles;Fluency in English;Excellent ...

Site Reliability Engineer (SRE) - Phoenix, AZ

Hiring Organisation
Ellofant
Location
Phoenix, Arizona, United States
Employment Type
Any
Salary
USD Annual
5. Conduct blameless post-mortems and root cause analysis to drive continuous improvement 6. Design and implement automated monitoring and alerting systems using Dynatrace, Grafana, Logscale, CrowdStrike, Prometheus, Splunk, Moogsoft, and Datadog 7. Create robust dashboards and implement SLAs/SLOs through comprehensive monitoring 8. Analyze metrics from operating systems … 1. 3-5 years of relevant experience in site reliability, infrastructure, or DevOps engineering 2. Strong expertise in monitoring and observability tools (Dynatrace, Grafana, Prometheus, Splunk, or similar) 3. Experience with incident management and event correlation platforms (BigPanda, ServiceNow, Moogsoft) 4. Proficiency with Linux/Unix systems (RHEL) and Windows ...

Engineer C# (Full Stack)

Location
City Of London, England, United Kingdom
Group Overview The TP ICAP Group is a world‐leading provider of market infrastructure. Our purpose is to provide clients with access to global financial and commodities markets, improving price discovery, liquidity, and distribution of ...