526 to 550 of 890 Grafana Jobs in the UK

Software Engineer, AI Libraries

Hiring Organisation
wayve
Location
London, United Kingdom
Salary
£ 80 K
working with large GPU clusters or distributed training environments.Familiarity with distributed training techniques such as DDP or FSDP.Experience with observability tools such as Prometheus, Grafana, Datadog, or OpenTelemetry.Experience with data pipeline orchestration tools such as Airflow, Flyte, Ray, Metaflow, or Argo Workflows.Experience with containerisation and infrastructure tooling such as Docker ...

Neo4j Platform Consultant

Hiring Organisation
NTT DATA
Location
London, United Kingdom
Salary
£ 80 K
Security (RBAC, authentication, data protection)Preferred SkillsExperience with Graph RAG workloads and traversal optimizationDevOps tools (Docker, Kubernetes, CI/CD pipelines)Monitoring tools (Prometheus, Grafana, etc.)Experience with other graph databases (Neptune, TigerGraph)Job ExpectationsEnsure stable, secure, and high-performing Neo4j platform operationsEnable efficient graph query execution ...

Platform Engineer

Location
Blantyre, Scotland, United Kingdom
experience across the monitoring stack: Metrics collection (Zabbix, Prometheus, Nagios or similar) Log collection and aggregation (Graylog, Loki, Logstash or similar) Visualisation and dashboarding (Grafana, Kibana or similar) Working with technologies you have never seen before — a track record of being handed something unfamiliar and getting to a working result ...

Lead Site Reliability Engineer

Hiring Organisation
JP Morgan Chase
Location
London, United Kingdom
Salary
£ 80 K
hands on experience in front office trading environments or similarly high pressure, low latency domains.Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ (IBM MQ or similar), Oracle DBDemonstrated experience using enterprise-authorized AI capabilities within the work environment to improve ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
experience in front office trading environments or similarly high pressure, low latency domains. Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ (IBM MQ or similar), Oracle DB Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve ...

Lead Site Reliability Engineer

Location
Westminster, West End, United Kingdom
experience in front office trading environments or similarly high pressure, low latency domains. Proficiency with SRE tooling and techniques, including FIX messaging, Kafka, Grafana, Splunk, ITRS Geneos, Dynatrace, InfluxDB, MQ (IBM MQ or similar), Oracle DB Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve ...

Lead Product Manager

Location
Greater London, England, United Kingdom
inform product roadmap Our Tech Stack Languages: Go · Python · Node.js Infrastructure: Google Cloud · Docker · Kubernetes · Rancher · Jenkins CI/CD Edge & CDN: Cloudflare Observability: Grafana · Metabase AI/ML: Anthropic Claude · Google Gemma · Cursor/Warp Equal opportunity declaration Verifymy is an equal opportunity employer and is committed to providing ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost efficiency for LLM inference workloads using performance and optimization ...

Lead SRE - Chase UK

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
service discovery, ingress, networking, and load balancing. Experience with Kubernetes. Experience with cloud computing services. Familiarity with common observability and reliability toolchains such as Grafana, Prometheus, Elasticsearch, Kibana, or Jaeger. Ability to use AI-assisted engineering tools responsibly, including validating outputs, understanding failure modes, and applying secure handling of sensitive ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Paisley, Scotland, United Kingdom
clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost efficiency for LLM inference workloads using performance and optimization ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Milton Keynes, England, United Kingdom
clusters, using reproducible infrastructure as code and continuous delivery pipelines Implement observability (logs, metrics, traces) with dashboards and actionable alerting, including Prometheus metrics and Grafana/Alertmanager integration for LLM and GPU workloads Tune GPU and accelerator capacity, autoscaling, and cost efficiency for LLM inference workloads using performance and optimization ...

Software Engineering Specialist

Hiring Organisation
BT Group
Location
Cheltenham, Gloucestershire, United Kingdom
Salary
£ 70 K
Elastic stack technologies and Kibana plugin developmentHave used DevOps tools and principles like Git, Jenkins, GitOpsUsed system monitoring tools such as CheckMK, Prometheus, Grafana or LokiOur PackageTailored benefits make a real difference. That’s why we offer a comprehensive range to support your growth, wellbeing, and everyday life. ...

AI Engineer

Hiring Organisation
AECOM
Location
London, UK
Employment Type
Full-time
Prompt engineering, Data processing Preferred Skills Understanding of optimizing both CPU-bound and GPU-bound workloads. Knowledge of monitoring and observability tools (e.g. Prometheus, Grafana, Elastic stack). Experience with Infrastructure as Code (Terraform). Experience with Azure Experience with Machine Learning Experience in building product for the construction industry. ...

Principal Machine Learning Infrastructure Engineer London, United Kingdom

Location
Greater London, England, United Kingdom
consume data Experience building model serving infrastructure with latency and throughput requirements Familiarity with experiment tracking tools (Weights & Biases, MLflow) and observability stacks (Prometheus, Grafana) What we offer Equity options – share in our success and growth. 10% employer pension contribution – invest in your future. Free office lunches – great food ...

Senior Java Software Engineer - Intelligent Operations

Location
Cardiff, Wales, United Kingdom
Distributed service based architecture Kubernetes (EKS) TeamCity for CI/CD (lots of teams are releasing code 15-20 times per day!) Terraform and Grafana Our process Interviewing is a two way process and we want you to have the time and opportunity to get to know us, as much ...

Principal Site Reliability Engineer

Hiring Organisation
Oracle Corporation
Location
United Kingdom
Salary
£ 60 K
Must support network segmentation (e.g., security lists, network security groups, or firewalls).-Deep Understanding of manipulating telemetry data (traffic flows, health status) using Grafana dashboards and MQL.-Experience with major public cloud providers (e.g., Oracle Cloud Infrastructure OCI, or equivalent).-Experience using Jira and Confluence for incident tracking ...

Java Software Engineer - Customer Services

Location
Greater London, England, United Kingdom
Distributed service based architecture Kubernetes (EKS) TeamCity for CI/CD (lots of teams are releasing code 15-20 times per day!) Terraform and Grafana The team The Customer Service Engineering group is all about making things better for our customers. We use the latest tech to build solid solutions ...

Platform Engineer (10x Openings)

Location
Greater London, England, United Kingdom
similar) at an engineering level. Background building Kubernetes operators using frameworks such as Kopf, controller‐runtime, or similar. Experience with observability tooling: Prometheus, Grafana, OpenTelemetry, or structured logging in distributed systems. Experience building SaaS or PaaS layers on top of an IaaS platform. Exposure to serverless or inference serving infrastructure. ...

Database Reliability Engineer

Hiring Organisation
Starling Bank
Location
London, United Kingdom
Salary
£ 80 K
cloud—while ensuring rigorous data integrity and mobilityA Security & Observability Mindset: You believe security is paramount. You focus on building deep observability (Prometheus/Grafana/OpenTelemetry/Humio) and automated guardrails so the fleet is secure by design without requiring manual interventionEngineering via Code: While you are a systems ...

Senior Infrastructure Engineer, Research Singapore

Location
Greater London, England, United Kingdom
consume data Experience building model serving infrastructure with latency and throughput requirements Familiarity with experiment tracking tools (Weights & Biases, MLflow) and observability stacks (Prometheus, Grafana) What we offer Build what actually matters Help shape an AI-native engineering company at a formative stage, tackling problems that genuinely matter for industry ...

IT Innovation Operations Specialist (AI) - IT Services - 107877 - Grade 7

Hiring Organisation
University of Birmingham
Location
United Kingdom
Salary
£ 70 K
platforms, SaaS applications, or complex software suites.Monitoring & Observability: Strong skills in setting up and interpreting operational monitoring views, dashboard systems (e.g., Azure Monitor, Grafana, AWS CloudWatch, or similar), and analyzing application logs to troubleshoot performance.Cloud Infrastructure Knowledge: Working operational knowledge of modern cloud environments such as Azure (Microsoft ...

Senior AWS Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£70,000
Stand Out If You Have: Practical experience managing large-scale Kubernetes clusters; certifications in Kubernetes are a strong bonus Hands-on familiarity with the Grafana Observability Suite, including tools like Loki, Mimir, and Tempo Background in administering or developing with popular monitoring and automation tools such as Splunk, Datadog, PagerDuty …/CD, or CircleCI Strong understanding of containerization (e.g., Docker, Kubernetes) and microservices architecture Skilled in using observability and monitoring tools such as Prometheus, Grafana, ELK stack, or AWS CloudWatch Excellent analytical and troubleshooting abilities, especially within complex distributed systems Proven experience handling incident management and conducting blameless postmortems, including ...

Senior Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment Limited
Location
Southampton, Hampshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£70,000
Stand Out If You Have: Practical experience managing large-scale Kubernetes clusters; certifications in Kubernetes are a strong bonus Hands-on familiarity with the Grafana Observability Suite, including tools like Loki, Mimir, and Tempo Background in administering or developing with popular monitoring and automation tools such as Splunk, Datadog, PagerDuty …/CD, or CircleCI Strong understanding of containerization (e.g., Docker, Kubernetes) and microservices architecture Skilled in using observability and monitoring tools such as Prometheus, Grafana, ELK stack, or AWS CloudWatch Excellent analytical and troubleshooting abilities, especially within complex distributed systems Proven experience handling incident management and conducting blameless postmortems, including ...

Site Reliability Engineer III

Location
Belfast City District, Northern Ireland, United Kingdom
Job Title: Site Reliability Engineer (SRE) III – Platform Engineering & Systems Reliability The Role: CME Group is seeking a Site Reliability Engineer (SRE) III to engineer reliability for our Google Cloud (GCP) infrastructure, Middleware Platform Engineering ...

Test Environment Manager

Hiring Organisation
Euroclear
Location
United Kingdom
Salary
£ 70 K
Job description:We are seeking an experienced Test Environment Manager to drive the test environment management to support a large-scale transformation programme moving from a legacy monolithic setup to a microservices-based set up ...