726 to 750 of 788 Permanent Grafana Jobs

Infrastructure Lead

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
An innovator in the artificial space is looking for a talented team lead to join their Infrastructure team. In this role, you will lead projects, and design and implement architectural changes to solve complex technical ...

Sr. Network Site Reliability Engineer (SREs)

Location
Greater London, England, United Kingdom
automation workflows using Ansible, Salt, and related frameworks to reduce operational toil. Build and operate monitoring, alerting, and observability dashboards using tools such as Grafana and Splunk. Proactively identify network bottlenecks, performance issues, and reliability risks, implementing long‐term fixes rather than reactive solutions. Support incident response, root cause analysis … OSPF, EIGRP, STP, VXLAN, VPNs, QoS, MPLS, etc.). Strong experience with infrastructure automation using Ansible and Salt. Proficiency with observability tooling such as Grafana, Splunk, or equivalent. Solid understanding of SRE practices including SLIs, SLOs, error budgets, and proactive reliability. Strong troubleshooting, analytical, and performance optimization skills. Excellent communication ...

Network Engineer

Location
Greater London, England, United Kingdom
security posture of network and network security platforms end to end Enhancing observability across telemetry, alerting and performance monitoring using tools such as Grafana, Prometheus and ThousandEyes to enable proactive operations Implementing and maintaining scalable, resilient datacentre network infrastructure built on Cisco and Arista technologies Championing automation of operational tasks … Network Engineer in enterprise or large-scale environments Experience applying SRE, observability and automation principles to networking, using technologies such as Python, Prometheus, Grafana, OpenTelemetry, Ansible and Jenkins Experience with Cisco and Arista switching and routing, alongside network security infrastructure, including firewalls, IDS/IPS and network segmentation Expertise ...

Cloud Support Engineer

Location
United Kingdom
knowledge of Microsoft Entra ID, RBAC, Conditional Access, PIM, Defender for Cloud, and Azure Policy. Experience with Azure Monitor, Log Analytics, Application Insights, DCRs, Grafana, and KQL. Infrastructure as Code experience with Terraform, Bicep, and ARM Templates. Strong PowerShell and Azure CLI scripting and automation skills. Experience with Azure DevOps … applications, and security. Implement Azure security, governance, compliance, and identity controls. Design and improve monitoring and alerting using Azure Monitor, Log Analytics, Application Insights, Grafana, and KQL. Develop Infrastructure as Code using Terraform, Bicep, and ARM Templates. Automate operational tasks using PowerShell, Azure CLI, and Power Automate. Support customer migrations ...

Production Support Engineer

Hiring Organisation
Pinnacle Technical Resources
Location
Charlotte, North Carolina, United States
Employment Type
Permanent
Salary
USD 50 Hourly
production support for critical applications, performing root cause analysis (RCA) and implementing long-term solutions. Monitor application and infrastructure performance using tools such as Grafana, AppDynamics, Splunk, and ITRS, and proactively resolve issues. Develop, maintain, and enhance automation scripts and services using Ansible, Python, Shell, PowerShell, and database queries … support role. Hands-on experience with OpenShift, Ansible, and automation scripting (Python, Shell, PowerShell, etc.). Proficiency in monitoring and observability tools such as Grafana, AppDynamics, ITRS, and Splunk. Strong knowledge of databases (SQL, queries, optimization). Experience supporting large-scale, high-availability systems in a 24x7 production environment. Familiarity ...

Senior Software Engineer- Python

Location
Manchester, England, United Kingdom
serverless architectures at meaningful scale Infrastructure as Code (CDK, CloudFormation, or Terraform) Experience with Pydantic, strict typing, and mypy Observability in production (CloudWatch, Grafana) Exposure to energy, trading, optimisation, or other constraint heavy domains Working in a monorepo with shared libraries and multiple deployable services You do not need … integrate with the solver) Infrastructure and tooling AWS CDK (TypeScript) and CloudFormation aws-lambda-powertools for observability CI/CD via AWS CodePipeline Grafana dashboards and CloudWatch alarms Product and analysis Databricks SQL, dashboards and notebooks for reporting Jupyter, numpy, pandas, plotly for investigation Cursor, Copilot, and a healthy amount ...

Senior Network Engineer

Location
Greater London, England, United Kingdom
Using BGP, EVPN-VXLAN, JunOS (Juniper QFX/MX), OSPF, Spine-Leaf/IP Fabric, VLANs, VRFs, Linux networking, Ansible, Terraform, Prometheus/Grafana, Observability tooling The adventures that await you after becoming Senior Network Engineer at Hack The Box: Design and implement spine-leaf network architectures across multiple data … JunOS Develop and maintain network automation using Ansible, Terraform or similar infrastructure‐as‐code tooling Establish and improve network monitoring, alerting and observability (Prometheus, Grafana, SNMP, streaming telemetry) Plan and execute network capacity upgrades, site bring‐ups and hardware refresh cycles Collaborate with the platform/systems engineering team ...

Senior Network Engineer, Studios

Location
Greater London, England, United Kingdom
through CI pipelines (Python, Ansible, Git); eliminate hand‐edits. Observability: develop monitoring built on streaming telemetry (gNMI/gRPC), flow analysis and modern tooling (Grafana, Zabbix class), serving the 24/7 operations teams as your internal customers. Modernisation: lead the shift from manual workflows to automated NetOps across … history. IPAM/DCIM: experience with NetBox or a similar platform for infrastructure documentation and management; NetBox preferred. Observability: practical experience with modern monitoring (Grafana, Zabbix, ELK class), flow analysis and streaming telemetry. Security: proven experience operating and securing enterprise firewall estates (Palo Alto, Fortinet or Cisco class) and managing ...

Senior Full Stack Engineer (Java + React)

Hiring Organisation
Luxoft
Location
London, UK
Employment Type
Full-time
Project descriptionJoin the Equity Accelerator programme, a strategic multi-year transformation initiative focused on modernizing Prime Finance Equities platforms. The role involves developing and enhancing institutional-grade trading and post-trade systems using cloud-native ...

Senior Full Stack Engineer (Kotlin + React)

Hiring Organisation
Luxoft
Location
London, UK
Employment Type
Full-time
Project descriptionJoin the Equity Accelerator programme, a strategic multi-year transformation initiative focused on modernizing Prime Finance Equities platforms. The role involves developing and enhancing institutional-grade trading and post-trade systems using cloud-native ...

Senior Full Stack Engineer (Kotlin + React)

Location
Greater London, England, United Kingdom
Project description Join the Equity Accelerator programme, a strategic multi-year transformation initiative focused on modernizing Prime Finance Equities platforms. The role involves developing and enhancing institutional-grade trading and post-trade systems using cloud ...

Senior Site Reliability Engineer (Observability)

Location
United Kingdom
engineering team better at running its own services. Today our telemetry lives in Azure Monitor, Log Analytics and Application Insights, with Azure Managed Grafana for dashboards and alerting. It works, but it has grown organically. Alert thresholds are inherited rather than designed, retention and cost are not governed, instrumentation … target architecture and standards for the platform - what we instrument, how, where it lands, how long we keep it and what it costs. Make Grafana the place engineers go to understand production: dashboards and alerting designed around services and customer journeys, not around whichever metrics happened to be available. Establish ...

Network Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
availability and security posture of network and network security platforms end to endEnhancing observability across telemetry, alerting and performance monitoring using tools such as Grafana, Prometheus and ThousandEyes to enable proactive operationsImplementing and maintaining scalable, resilient datacentre network infrastructure built on Cisco and Arista technologiesChampioning automation of operational tasks using … Network Engineer in enterprise or large-scale environmentsExperience applying SRE, observability and automation principles to networking, using technologies such as Python, Prometheus, Grafana, OpenTelemetry, Ansible and JenkinsExperience with Cisco and Arista switching and routing, alongside network security infrastructure, including firewalls, IDS/IPS and network segmentationExpertise in Cisco and Arista ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent, Work From Home
understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. Proficiency in shell scripting for automation and system management. Experience with Infrastructure as Code, including Terraform and Ansible. Knowledge of Cloudflare … operational consistency. Write and contribute to code, telemetry and instrumentation that improve service reliability and observability. Build dashboards and operational views using telemetry from Grafana, Splunk, New Relic and related platforms. Configure and manage Cloudflare edge services using Infrastructure as Code and integrate edge telemetry with observability platforms. Diagnose incidents ...

Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Stoke-On-Trent, Staffordshire, West Midlands, United Kingdom
Employment Type
Permanent, Work From Home
understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management. Hands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty. Proficiency in shell scripting for automation and system management. Experience with Infrastructure as Code, including Terraform and Ansible. Knowledge of Cloudflare … operational consistency. Write and contribute to code, telemetry and instrumentation that improve service reliability and observability. Build dashboards and operational views using telemetry from Grafana, Splunk, New Relic and related platforms. Configure and manage Cloudflare edge services using Infrastructure as Code and integrate edge telemetry with observability platforms. Diagnose incidents ...

Senior Data Engineer

Location
Greater London, England, United Kingdom
Staff level. Embed DataOps best practices across the team: CI/CD, testing, observability, data drift monitoring, and incident response using tools like Grafana and Datadog. Collaborate with ML engineers and data scientists to productionise models and agentic tooling built on internal AI platforms, with hands … build expertise in stream processing technologies such as Apache Flink, Kafka, or Spark. (Nice to have: ClickHouse.) Familiarity with observability tooling such as Grafana or Datadog, and a good instinct for keeping systems healthy and well-monitored. Active engagement with the evolving AI/ML tooling landscape — you integrate ...

Senior Backend Engineer - Asset Sales

Location
Greater London, England, United Kingdom
C# stack : Distributed C# and .NET microservices Cloud & orchestration : Hosted on Azure using Kubernetes Architecture : Event-driven, supporting products used at significant scale Observability : Grafana, Azure Application Insights, logs, traces, and metrics AI tooling : Claude and other AI tools used throughout the engineering workflow — design exploration, code generation and review … have experience with similar messaging technology (Kafka, RabbitMQ, etc.) Microservices architecture experience, ideally on Azure and Kubernetes Experience using observability tooling (e.g. Grafana, Application Insights) to understand production behaviour Comfort using AI tools as part of a daily engineering workflow (design, code review, testing, incident investigation) DevOps culture mindset, including ...

Database Reliability Engineer

Location
Manchester, England, United Kingdom
Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will leverage Java-based data collection pipelines and metrics platforms (Prometheus, Grafana, Dash0, Sentry, Humio) to detect performance regressions, slow queries, and health issues before they impact customers Fortify Business Continuity (BCP): Design and implement rigorous Business … region data platforms—while ensuring data integrity, clean relational modeling, and mobility A Security & Observability Mindset: You focus on building deep observability (Prometheus/Grafana/Dash0/Sentry/Humio) and automated security guardrails directly into your Java services so the fleet is secure and monitored by design Interview ...

Database Reliability Engineer

Location
Cardiff, Wales, United Kingdom
Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will leverage Java-based data collection pipelines and metrics platforms (Prometheus, Grafana, Dash0, Sentry, Humio) to detect performance regressions, slow queries, and health issues before they impact customers Fortify Business Continuity (BCP): Design and implement rigorous Business … region data platforms—while ensuring data integrity, clean relational modeling, and mobility A Security & Observability Mindset: You focus on building deep observability (Prometheus/Grafana/Dash0/Sentry/Humio) and automated security guardrails directly into your Java services so the fleet is secure and monitored by design Interview ...

Database Reliability Engineer

Location
Southampton, England, United Kingdom
Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will leverage Java-based data collection pipelines and metrics platforms (Prometheus, Grafana, Dash0, Sentry, Humio) to detect performance regressions, slow queries, and health issues before they impact customers Fortify Business Continuity (BCP): Design and implement rigorous Business … region data platforms—while ensuring data integrity, clean relational modeling, and mobility A Security & Observability Mindset: You focus on building deep observability (Prometheus/Grafana/Dash0/Sentry/Humio) and automated security guardrails directly into your Java services so the fleet is secure and monitored by design Interview ...

Database Reliability Engineer

Location
Greater London, England, United Kingdom
Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will leverage Java-based data collection pipelines and metrics platforms (Prometheus, Grafana, Dash0, Sentry, Humio) to detect performance regressions, slow queries, and health issues before they impact customers Fortify Business Continuity (BCP): Design and implement rigorous Business … region data platforms—while ensuring data integrity, clean relational modeling, and mobility A Security & Observability Mindset: You focus on building deep observability (Prometheus/Grafana/Dash0/Sentry/Humio) and automated security guardrails directly into your Java services so the fleet is secure and monitored by design Interview ...

Senior Java Developer

Hiring Organisation
Luxoft
Location
London, UK
Employment Type
Full-time
Project descriptionJoin the Equity Accelerator programme to modernize critical Prime Finance Equities trading platforms. The role focuses on building and evolving highly scalable backend systems using Kotlin and Java, with event-driven, cloud-native architectures. ...

Technical Product Manager (Superapp)

Location
Greater London, England, United Kingdom
About LendableLendable is on a mission to build the world's best technology to help people get credit and save money. We're building one of the world’s leading fintech companies and are off ...

Solutions Integration Engineer

Hiring Organisation
Christy Media Solutions
Location
London, UK
Employment Type
Full-time
Reference: JOB-8544Christy Media Solutions is recruiting for a newly created Solutions Integration (DevOps) engineer position supporting our client's transition towards IP-centric, Software-as-a-Service broadcast services. Working within the architecture function ...

Senior Kotlin Developer

Hiring Organisation
Luxoft
Location
London, UK
Employment Type
Full-time
Project descriptionJoin the Equity Accelerator programme to modernize critical Prime Finance Equities trading platforms. The role focuses on building and evolving highly scalable backend systems using Kotlin and Java, with event-driven, cloud-native architectures. ...