676 to 700 of 962 Grafana Jobs

DevOps Engineer

Hiring Organisation
Sanderson Recruitment
Location
Romsey, Hampshire, South East, United Kingdom
Employment Type
Contract
Contract Rate
Up to £550 per day + Outside IR-35
hands-on experience with a modern DevOps toolchain, including: AWS Docker GitLab CI/CD JFrog Artifactory Kubernetes Helm Terraform ArgoCD Python Rancher Packer Grafana Prometheus OpenTelemetry Conventional Commits The Role Working within a highly technical engineering environment, you'll be responsible for building, automating and maintaining cloud infrastructure … optimisation Kubernetes platform administration and automation Infrastructure as Code using Terraform and Packer GitOps implementation with ArgoCD Monitoring, observability and performance optimisation using Grafana, Prometheus and OpenTelemetry Continuous improvement of engineering standards, automation and release processes Requirements Active SC or DV Clearance UK-based Able to attend site in Romsey ...

Platform Engineer

Location
Greater London, England, United Kingdom
Infrastructure as Code Develop and maintain GitOps CI/CD pipelines Manage Kubernetes networking, service mesh and gateway technologies Improve platform observability using Grafana, Prometheus and OpenTelemetry Maintain platform security, resilience and automation Troubleshoot production platform issues and drive continuous improvement Work closely with architects to turn high-level designs … essential) Kubernetes platform engineering within production environments Terraform and Infrastructure as Code Docker, GitOps and CI/CD pipelines Linux and Bash scripting Grafana, Prometheus and OpenTelemetry Kubernetes networking and service mesh technologies Production platform operations, troubleshooting and automation You'll also be able to demonstrate: Experience owning technical implementation ...

Vice President, Site Reliability Engineering

Hiring Organisation
The Bank of New York Mellon
Location
London, UK
Employment Type
Full-time
Objectives, and service health measures aligned to operational and business priorities. Build and optimize monitoring, observability, and alerting capabilities using tools such as Prometheus, Grafana, AppDynamics, and Splunk. Apply AIOps capabilities to improve event correlation, anomaly detection, root cause analysis, predictive insights, and proactive issue prevention. Partner with engineering, infrastructure … Demonstrated ability to define and operationalize SLIs, SLOs, dashboards, alerts, and health indicators. Hands-on experience with enterprise monitoring and observability platforms including Prometheus, Grafana, AppDynamics, and Splunk. Strong troubleshooting, analytical, and problem-solving skills in complex distributed or production environments. Strong verbal and written communication skills, with the ability ...

Sr. Manager, Site Reliability

Location
Manchester, England, United Kingdom
across Engineering and Support. Select and stand up the primary observability platform, preferring extension of existing Omnicell contracts (DataDog, IBM/Instana, Prometheus/Grafana, OpenTelemetry, or other tooling already in use) over net-new procurement. Define the instrumentation standards all new services must meet. Partner with … knowledge of Docker, Helm, and Service Mesh technologies (Istio, Linkerd). Hands-on experience designing modern observability platforms using tools such as DataDog, Prometheus, Grafana, OpenTelemetry, Elasticsearch/Kibana, or equivalent — with an opinion about what a good telemetry stack looks like. Familiarity with integrating AI/ML-based anomaly ...

Cloud Platform Engineer

Hiring Organisation
Credence
Location
Mc Lean, Virginia, United States
Employment Type
Permanent
Salary
USD 180,000 Annual
At Credence, we support our clients' mission-critical needs, powered by technology. We provide cutting-edge solutions, including AI/ML, enterprise modernization, and advanced intelligence capabilities, to the largest defense and health federal organizations. ...

Lead SRE- Azure & GCP

Location
Glasgow, Scotland, United Kingdom
We have a Lead Site Reliability Engineer (SRE) opportunity within our Google Cloud Site Reliability Engineering team. As a Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platform - Cloud Foundational Services SRE organization ...

AVP, Observability & SRE Engineer

Location
Greater London, England, United Kingdom
seeking a Site Reliability Engineer - Assistant Vice President in London to drive end‐to‐end observability, migrate legacy monitoring to Google Cloud Observability and Grafana, and implement OpenTelemetry instrumentation across OpenShift/Kubernetes environments. The role emphasizes hands‐on deployment, automation (Ansible/Terraform), and collaboration with application teams ...

Hybrid Linux Automation Engineer – Travel Expensed

Location
Milton, Scotland, United Kingdom
scale enterprise infrastructure, including Linux, VMware, and F5 environments. You will build automation to accelerate patching and changes, strengthen validation, improve observability with Prometheus, Grafana, and Airflow, and work with Python and Ansible within a collaborative engineering team. #J-18808-Ljbffr ...

Senior Cloud Platform Engineer – Multi-Cloud & Observability

Location
Greater London, England, United Kingdom
load for developers and clients at scale. The role focuses on building resilient, scalable tooling with Python or Golang and integrating Prometheus, OpenTelemetry, and Grafana across the stack. Hybrid London-based works environment offered. #J-18808-Ljbffr ...

DevOps Engineer – Prop Trading

Hiring Organisation
Quant Capital
Location
London, United Kingdom
Salary
£ 70 K
responsible for designing and supporting highly available systems across a diverse technology stack.My client leverages the latest technologies such as Dock, Kubernetes, Prometheus and Grafana as well as CI/CD to meet the growing demands of the business.You will join a small global team who are focused ...

Senior Backend Engineer

Location
Greater London, England, United Kingdom
pipelines with their backend core. Tech stack you will be working with: Go, Python, Postgres, Redis, NATS, Keycloak, Kubernetes, Istio, mTLS, ArgoCD, Prometheus, Grafana, AWS, Terraform. Key skills/experience required: Experience building and operating scalable production systems Experience with Go and Python, alongside as many areas of the tech ...

Network Operations Engineer

Location
Greater London, England, United Kingdom
networks, or comparable infrastructure Scripting proficiency in Python or Bash with hands-on Linux experience Practical experience with monitoring and visualization tools such as Grafana or Prometheus Strong communication skills and ability to perform effectively under pressure Nice to Have Background in satellite operations or satellite laser ranging systems Networking ...

Senior Data Engineer

Location
Greater London, England, United Kingdom
Experience building efficient, scalable databases and APIs (e.g. Django, FastAPI) a huge plus; Experience in Kubernetes, Docker, Spark and related monitoring tools (e.g. DataDog, Grafana, Prometheus) for DataOps a huge plus; Experience with Airflow a huge plus; Experience with dbt for pipeline modelling also beneficial; Skilled at shaping needs into ...

Machine Learning Engineer

Hiring Organisation
Understanding Recruitment
Location
Stevenage, Hertfordshire, United Kingdom
Salary
£ 60 K
maintaining production ML or AI servicesGood understanding of CI/CD, containers and infrastructure-as-codeExperience with monitoring and observability tools such as Grafana, CloudWatch, Langfuse or similarSome practical exposure to LLMs, RAG, NLP or generative AIThe confidence to take ownership of production systems and help guide other engineers ...

Site Reliability Engineer

Location
West of England, England, United Kingdom
production infrastructure operations, together with strong hands-on Python automation skills. You'll also need experience with: Monitoring and observability tools such as Prometheus, Grafana or similar Production incident management and/or on-call environments Automating operational runbooks and repetitive infrastructure processes APIs and systems integration Version-controlled automation ...

Software Developer - Data Reliability

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
vendorsMindset: Proactive, detail-oriented, and self-driven with a strong sense of ownership and accountabilityNice to haveExperience with observability and monitoring tools such as Grafana, Kibana, or PrometheusExperience developing automation tooling and implementing configuration managementExperience with cloud platforms such as Google Cloud or AWSExperience operating job orchestration or workload scheduling ...

Senior Performance Tester - SC Eligible/cleared

Hiring Organisation
scrumconnect ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 45,000 - 50,000 Annual
hybrid delivery models Experience working in regulated or high-assurance environments Desirable Experience with APM tools (eg Dynatrace, AppDynamics, Elastic Search with Kibana/Grafana across digital services ) Public-sector or large programme delivery experience Familiarity with GDS Scope & Accountability Responsible for hands-on performance test design and execution Owns ...

Site Reliability Engineer, Big Data (Remote, International)

Location
United Kingdom
Kafka for messaging layer Hadoop and Ceph as distributed storage layer SQL Server backup and recovery Terraform, Ansible, Puppet, ArgoCD for operational automation Prometheus, Grafana, Icinga and PagerDuty as observability layer Bare-metal servers and hybrid cloud/on-prem infrastructure What matters: understand distributed systems, failure recovery, and operational ...

Staff Backend Engineer

Location
Greater London, England, United Kingdom
modern, high-performance engineering stack designed for scale and reliability: Java 21 (Spring), Go, Node.js (TypeScript), PostgreSQL, MariaDB, Redis, Kafka, ClickHouse, Elasticsearch, Docker, Grafana, AWS, Kubernetes and Kibana. About the role Staff Backend Engineers shape technical direction across teams while remaining close to the systems and code. You will take ...

Senior UI Platform Engineer, London

Location
Greater London, England, United Kingdom
both· Experience with UI/component libraries such as Ant Design, AG Grid, or similar· Experience with observability tooling such as Datadog, OpenTelemetry, or Grafana· Experience with Microsoft Entra ID/Azure AD or similar identity platforms· Experience publishing and maintaining internal npm packages· Experience building admin, control-plane ...

Platform Engineer

Location
United Kingdom
distributed systems. Understanding of AI/ML concepts and tooling (model training, inference, data pipelines basics). Experience with monitoring/logging tools (e.g., Grafana, Kibana, Splunk). Excellent communication skills to interface with both customers and internal/vendor teams. Good understanding of tools requirements for ML engineers ...

EU Parnter Tech Lead

Hiring Organisation
ISG Personalmanagement
Location
United Kingdom
Salary
£ 60 K
interoperability.Exposure to cloud-native telecommunications environments.Technologies LTE • 5G NR • eNodeB • gNodeB • 3GPP • PHY • MAC • RLC • RRC • MIMO • Beamforming • HARQ • ICIC • CPRI • eCPRI • O-RAN • Grafana • Kibana • Power BI • Jira • ServiceNow • Confluence • Python • Bash • XCAL • XCAP • QXDM Equal Opportunity Statement We are committed to creating an inclusive workplace where diversity ...

AI Support Analyst

Location
Churwell, England, United Kingdom
analytical and problem-solving skills, with the ability to interpret complex, high-volume system data Familiarity with observability and monitoring tools such as Datadog, Grafana, Prometheus, Amplitude or similar Confidence working with AI-assisted tools and the ability to extract insight through effective prompting Excellent written and verbal communication skills ...

NOC Engineer / SRE

Location
United Kingdom
health monitoring across infrastructure, applications, and dependencies Design and maintain alerting strategies that align with SLIs/SLOs Build dashboards using tools such as: Grafana Reliability Engineering & Automation Automate repetitive operational tasks to reduce manual toil Improve mean time to detect (MTTD) and mean time to resolve (MTTR) Develop scripts ...