676 to 700 of 969 Grafana Jobs

Senior QA Engineer Performance & Scalability

Location
Greater London, England, United Kingdom
real, hands‐on depth with k6 for building and running performance tests. Monitoring and observability know‐how. You're comfortable with tools like Datadog, Grafana, CloudWatch or similar, and you use them proactively to correlate test results with what the system is actually doing, not just when something is already ...

Service Reliability Support Manager

Location
Greater London, England, United Kingdom
environment, including experience redesigning a reactive, manual function into an engineering-led operating model. Hands-on experience with observability tools and stacks such as Grafana, Prometheus, Open Telemetry, ELK, Splunk, or similar platforms. Deep understanding of SLIs, SLOs, error budgets, and telemetry best practices in high-availability environments. Excellent understanding ...

Sr. Software Engineer, Inference

Location
Greater London, England, United Kingdom
networked systems and performance optimisation. Hands-on experience with Kubernetes at production scale, including automated CI/CD and modern observability stacks (e.g., Prometheus, Grafana, OpenTelemetry). Practical, working knowledge of inference internals: batching strategies, caching, mixed precision (BF16/FP8), and streaming token delivery. Proven track record of improving ...

Senior DevOps Engineer

Hiring Organisation
Genesys
Location
Portola Valley, California, United States
Employment Type
Any
Salary
USD 103,000 Annual
Build and maintain CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI/CD, and Argo CD. Implement monitoring and observability using Prometheus, Grafana, CloudWatch, Datadog, Splunk, and PagerDuty. Develop automation scripts using Python, Bash, and PowerShell. Support production systems, incident response, root-cause analysis, disaster recovery, and high … processes, platform reliability, scalability, security, and developer productivity. Required Skills AWS, Kubernetes, EKS, Docker, OpenShift, Terraform, Ansible, Jenkins, GitHub Actions, Argo CD, Helm, Prometheus, Grafana, CloudWatch, Python, Bash, Linux, GitOps, CI/CD, SRE, and Infrastructure as Code. Preferred Qualifications Bachelor's or Master's degree in Computer Science, Information ...

DevOps Engineer

Hiring Organisation
Sanderson Recruitment
Location
Romsey, Hampshire, South East, United Kingdom
Employment Type
Contract
Contract Rate
Up to £550 per day + Outside IR-35
hands-on experience with a modern DevOps toolchain, including: AWS Docker GitLab CI/CD JFrog Artifactory Kubernetes Helm Terraform ArgoCD Python Rancher Packer Grafana Prometheus OpenTelemetry Conventional Commits The Role Working within a highly technical engineering environment, you'll be responsible for building, automating and maintaining cloud infrastructure … optimisation Kubernetes platform administration and automation Infrastructure as Code using Terraform and Packer GitOps implementation with ArgoCD Monitoring, observability and performance optimisation using Grafana, Prometheus and OpenTelemetry Continuous improvement of engineering standards, automation and release processes Requirements Active SC or DV Clearance UK-based Able to attend site in Romsey ...

Platform Engineer

Location
Greater London, England, United Kingdom
Infrastructure as Code Develop and maintain GitOps CI/CD pipelines Manage Kubernetes networking, service mesh and gateway technologies Improve platform observability using Grafana, Prometheus and OpenTelemetry Maintain platform security, resilience and automation Troubleshoot production platform issues and drive continuous improvement Work closely with architects to turn high-level designs … essential) Kubernetes platform engineering within production environments Terraform and Infrastructure as Code Docker, GitOps and CI/CD pipelines Linux and Bash scripting Grafana, Prometheus and OpenTelemetry Kubernetes networking and service mesh technologies Production platform operations, troubleshooting and automation You'll also be able to demonstrate: Experience owning technical implementation ...

Vice President, Site Reliability Engineering

Hiring Organisation
The Bank of New York Mellon
Location
London, UK
Employment Type
Full-time
Objectives, and service health measures aligned to operational and business priorities. Build and optimize monitoring, observability, and alerting capabilities using tools such as Prometheus, Grafana, AppDynamics, and Splunk. Apply AIOps capabilities to improve event correlation, anomaly detection, root cause analysis, predictive insights, and proactive issue prevention. Partner with engineering, infrastructure … Demonstrated ability to define and operationalize SLIs, SLOs, dashboards, alerts, and health indicators. Hands-on experience with enterprise monitoring and observability platforms including Prometheus, Grafana, AppDynamics, and Splunk. Strong troubleshooting, analytical, and problem-solving skills in complex distributed or production environments. Strong verbal and written communication skills, with the ability ...

Sr. Manager, Site Reliability

Location
Manchester, England, United Kingdom
across Engineering and Support. Select and stand up the primary observability platform, preferring extension of existing Omnicell contracts (DataDog, IBM/Instana, Prometheus/Grafana, OpenTelemetry, or other tooling already in use) over net-new procurement. Define the instrumentation standards all new services must meet. Partner with … knowledge of Docker, Helm, and Service Mesh technologies (Istio, Linkerd). Hands-on experience designing modern observability platforms using tools such as DataDog, Prometheus, Grafana, OpenTelemetry, Elasticsearch/Kibana, or equivalent — with an opinion about what a good telemetry stack looks like. Familiarity with integrating AI/ML-based anomaly ...

Cloud Platform Engineer

Hiring Organisation
Credence
Location
Mc Lean, Virginia, United States
Employment Type
Permanent
Salary
USD 180,000 Annual
At Credence, we support our clients' mission-critical needs, powered by technology. We provide cutting-edge solutions, including AI/ML, enterprise modernization, and advanced intelligence capabilities, to the largest defense and health federal organizations. ...

Lead SRE- Azure & GCP

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
We have a Lead Site Reliability Engineer (SRE) opportunity within our Google Cloud Site Reliability Engineering team. As a Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platform - Cloud Foundational Services SRE organization ...

Lead SRE- Azure & GCP

Location
Glasgow, Scotland, United Kingdom
We have a Lead Site Reliability Engineer (SRE) opportunity within our Google Cloud Site Reliability Engineering team. As a Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platform - Cloud Foundational Services SRE organization ...

AVP, Observability & SRE Engineer

Location
Greater London, England, United Kingdom
seeking a Site Reliability Engineer - Assistant Vice President in London to drive end‐to‐end observability, migrate legacy monitoring to Google Cloud Observability and Grafana, and implement OpenTelemetry instrumentation across OpenShift/Kubernetes environments. The role emphasizes hands‐on deployment, automation (Ansible/Terraform), and collaboration with application teams ...

Hybrid Linux Automation Engineer – Travel Expensed

Location
Milton, Scotland, United Kingdom
scale enterprise infrastructure, including Linux, VMware, and F5 environments. You will build automation to accelerate patching and changes, strengthen validation, improve observability with Prometheus, Grafana, and Airflow, and work with Python and Ansible within a collaborative engineering team. #J-18808-Ljbffr ...

Senior Cloud Platform Engineer – Multi-Cloud & Observability

Location
Greater London, England, United Kingdom
load for developers and clients at scale. The role focuses on building resilient, scalable tooling with Python or Golang and integrating Prometheus, OpenTelemetry, and Grafana across the stack. Hybrid London-based works environment offered. #J-18808-Ljbffr ...

SC Cleared Non-Functional Test Lead

Hiring Organisation
VIQU Limited
Location
London, UK
Employment Type
Full-time
ensuring that testing aligns with the performance strategy and agreed non-functional requirements. Working closely with technical teams to troubleshoot issues, monitor systems with Grafana and Splunk, and use diagnostics tools such as New Relic and Dynatrace. Later, you'll integrate tests into CI/CD pipelines and drive improvements ...

DevOps Engineer – Prop Trading

Hiring Organisation
Quant Capital
Location
London, United Kingdom
Salary
£ 70 K
responsible for designing and supporting highly available systems across a diverse technology stack.My client leverages the latest technologies such as Dock, Kubernetes, Prometheus and Grafana as well as CI/CD to meet the growing demands of the business.You will join a small global team who are focused ...

Senior Backend Engineer

Location
Greater London, England, United Kingdom
pipelines with their backend core. Tech stack you will be working with: Go, Python, Postgres, Redis, NATS, Keycloak, Kubernetes, Istio, mTLS, ArgoCD, Prometheus, Grafana, AWS, Terraform. Key skills/experience required: Experience building and operating scalable production systems Experience with Go and Python, alongside as many areas of the tech ...

Network Operations Engineer

Location
Greater London, England, United Kingdom
networks, or comparable infrastructure Scripting proficiency in Python or Bash with hands-on Linux experience Practical experience with monitoring and visualization tools such as Grafana or Prometheus Strong communication skills and ability to perform effectively under pressure Nice to Have Background in satellite operations or satellite laser ranging systems Networking ...

Senior Data Engineer

Location
Greater London, England, United Kingdom
Experience building efficient, scalable databases and APIs (e.g. Django, FastAPI) a huge plus; Experience in Kubernetes, Docker, Spark and related monitoring tools (e.g. DataDog, Grafana, Prometheus) for DataOps a huge plus; Experience with Airflow a huge plus; Experience with dbt for pipeline modelling also beneficial; Skilled at shaping needs into ...

Machine Learning Engineer

Hiring Organisation
Understanding Recruitment
Location
Stevenage, Hertfordshire, United Kingdom
Salary
£ 60 K
maintaining production ML or AI servicesGood understanding of CI/CD, containers and infrastructure-as-codeExperience with monitoring and observability tools such as Grafana, CloudWatch, Langfuse or similarSome practical exposure to LLMs, RAG, NLP or generative AIThe confidence to take ownership of production systems and help guide other engineers ...

Site Reliability Engineer

Location
West of England, England, United Kingdom
production infrastructure operations, together with strong hands-on Python automation skills. You'll also need experience with: Monitoring and observability tools such as Prometheus, Grafana or similar Production incident management and/or on-call environments Automating operational runbooks and repetitive infrastructure processes APIs and systems integration Version-controlled automation ...

Software Developer - Data Reliability

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
vendorsMindset: Proactive, detail-oriented, and self-driven with a strong sense of ownership and accountabilityNice to haveExperience with observability and monitoring tools such as Grafana, Kibana, or PrometheusExperience developing automation tooling and implementing configuration managementExperience with cloud platforms such as Google Cloud or AWSExperience operating job orchestration or workload scheduling ...

Senior Performance Tester - SC Eligible/cleared

Hiring Organisation
scrumconnect ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP 45,000 - 50,000 Annual
hybrid delivery models Experience working in regulated or high-assurance environments Desirable Experience with APM tools (eg Dynatrace, AppDynamics, Elastic Search with Kibana/Grafana across digital services ) Public-sector or large programme delivery experience Familiarity with GDS Scope & Accountability Responsible for hands-on performance test design and execution Owns ...

Site Reliability Engineer, Big Data (Remote, International)

Location
United Kingdom
Kafka for messaging layer Hadoop and Ceph as distributed storage layer SQL Server backup and recovery Terraform, Ansible, Puppet, ArgoCD for operational automation Prometheus, Grafana, Icinga and PagerDuty as observability layer Bare-metal servers and hybrid cloud/on-prem infrastructure What matters: understand distributed systems, failure recovery, and operational ...