476 to 500 of 526 Prometheus Jobs in England

Lead Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment Limited
Location
Southampton, Hampshire, South East, United Kingdom
Employment Type
Permanent
Salary
£85,000
level indicators and error budgets. Design and implement monitoring, alerting and dashboarding across cloud platforms and microservices. Deploy and configure observability technologies including Grafana, Prometheus, Azure Monitor and OpenTelemetry. Develop custom application and platform metrics to improve operational visibility. Create advanced queries, dashboards and alerts for distributed microservices. Develop reusable … using AKS. Extensive experience in platform engineering, cloud provisioning and observability. Strong monitoring, alerting and dashboarding experience using technologies such as: Azure Monitor, Grafana, Prometheus, OpenTelemetry, Elasticsearch Experience creating custom metrics, queries, dashboards and alerts for microservices. Advanced scripting or software development skills using PowerShell, Python, C# or a comparable ...

Cloud Architect - Kubernetes & Containerization Specialist

Location
Greater London, England, United Kingdom
LECP Head of Architecture Set up proactive monitoring, logging and incident response for all containerized applications using tools such as CloudWatch, ELK Stack and Prometheus to meet SLA requirements Work model At Cognizant, we strive to provide flexibility wherever possible. Based on this role's business requirements, this … controls, including GDPR, ISO27001 and NCSC guidelines for secure cloud deployments Knowledge of monitoring and logging solutions suitable for high-security environments, including CloudWatch, Prometheus, Grafana and centralized logging with ELK Strong analytical and problem-solving abilities, with a focus on creating resilient, secure and scalable solutions Excellent communication skills ...

DevOps and Automation Engineer (Contract)

Location
Greater London, England, United Kingdom
towards self-service environment provisioning using Infrastructure-as-Code (Terraform) and pipeline-driven automation.* Implement advanced observability and monitoring: Use platforms such as Datadog, Prometheus, Grafana, and OpenTelemetry to provide real-time insights into system health, deployments, and business metrics.* Embed security and compliance by design: Integrate security into every … generate new automation ideas and create user stories for rapid prototyping.* Use technologies like Terraform, Ansible, Azure DevOps, Github, OctopusDeploy, Kubernetes, OpenTelemetry, Datadog, Grafana, Prometheus, low-code automation platforms (e.g., Power Automate, UiPath) to help evolve the team's capabilities.* Coordinate planned outage and environment refreshes in collaboration with project ...

Lead Site Reliability Engineer

Location
Southampton, England, United Kingdom
Develop and configure monitoring dashboards and alerts in tools like Grafana and Azure Monitor. Installation and configuration of Observability Platform including tools like Grafana, Prometheus, Azure Monitor, Open telemetry etc. Developing bicep modules for monitoring infrastructure and deploy it. Optimize system performance, cost, and security through regular reviews and tuning. … etc.) Experience with infrastructure/configuration as code and version control (ARM, BICEP, Git) Strong Experience managing monitoring, alerting and dashboarding platforms (Azure Monitor, Prometheus, Grafana, Elasticsearch) Demonstrable experience of supporting live cloud services and platforms Expert in developing queries for dashboards and alerting for microservices. Collaborate with DevOps ...

Software Engineer II

Hiring Organisation
Couchbase
Location
Manchester, Greater Manchester, United Kingdom
Salary
£ 70 K
FeaturesWorking with senior engineers and product managers, design, implement, test and ship well-defined features that integrate Couchbase with Cloud Native tech such as Prometheus, Fluentd and Fluent-bit, and the Kubernetes Pod Autoscaler.Write clean, maintainable, well-tested code; participate in code reviews, both giving and receiving feedback. Communicate progress … test coverage for the features you build.Support Across LifecycleContribute to documentation, examples and tutorials on integration with Cloud Native ecosystem components such as Fluentd, Prometheus and OpenTelemetry.Help Sales Engineers, Professional Services and Support investigate and resolve customer issues related to the features you own.Participate in the team’s development processes ...

IT & Service Delivery Engineer

Location
Coleshill CP, England, United Kingdom
using Bash, Python or PerlTCP/IP networking, DNS, routing, VPNs, BGP, SMTP, HTTP and HTTPSPalo Alto firewalls; HAProxy, NGINX and dynamic DNS servicesGrafana, Prometheus, Graphite (ClickHouse), InfluxDB and Icinga, with alerting into OpsGenieELK for application logging and Wazuh for platform host security monitoringGit, Bitbucket, Jira and Bamboo CI/… with the confidence to deal with colleagues at any level.Automation and monitoringConfiguration management with Ansible, Puppet or AWX.Monitoring, alerting and dashboarding with Icinga, Grafana, Prometheus or similar.Infrastructure-as-code, CI/CD tooling or automated testing for infrastructure changes.AI-assisted coding and automation tooling used to accelerate operational and scripting ...

SC Cleared AWS DevOps & Platform Engineer (Remote)

Location
England, United Kingdom
iO Associates is seeking an SC Cleared AWS DevOps/Platform Engineer on a 6-month contract. The role is predominantly remote with occasional client site visits in Hampshire. You will design, deploy and operate ...

DevOps Engineer - Newcastle

Location
Newcastle upon Tyne, England, United Kingdom
Overview DevOps Engineer – Location: Newcastle Upon Tyne. Please note: due to the nature of client work, you will be required to undergo a Security Clearance process, which requires 5+ years UK address history at the ...

Senior Performance Engineer

Location
Cambridge, England, United Kingdom
models execute on that hardware (inference vs. training, matrix multiplication, KV-caching, etc.) Proficiency with profiling tools (NVIDIA Nsight, PyTorch Profiler) and monitoring stacks (Prometheus, Grafana) Capability to work in Python for data analysis (Pandas, NumPy) and scripting The following are also highly valued: Post-graduate degrees and research experience ...

Senior Performance Engineer | AI Infrastructure | Cambridge (Hybrid) |

Hiring Organisation
Pure Resourcing Solutions
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£90,000 - £120,000 per annum
training versus inference, matrix multiplication, KV-caching, that level of detail Comfortable in profiling tools like Nsight or PyTorch Profiler, and monitoring stacks like Prometheus and Grafana Python for data work, Pandas and NumPy, plus general scripting Nice to have rather than essential: a postgraduate degree and research background (publications ...

Performance Engineer (Junior) | AI Infrastructure | Cambridge (Hybrid)

Hiring Organisation
Pure Resourcing Solutions Limited
Location
Linton, Dry Drayton, Cambridgeshire, United Kingdom
Employment Type
Permanent
Salary
£55000 - £70000/annum
work with GPU or accelerator code, CUDA or similar Familiarity with profiling tools (Nsight, PyTorch Profiler) and ideally some exposure to monitoring stacks (Prometheus, Grafana) Strong Python for data work, Pandas and NumPy, genuine scripting ability Nice to have: exposure to inference serving frameworks like vLLM, published research, or open ...

Senior Network Engineer - Low latency trading

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Linux, covering boot sequence, service management, file, process, network, and kernel management. Ability to deploy and operate monitoring frameworks such as Cacti, Zabbix, Nagios, Prometheus, or Tick stack. Strong firewall experience (Checkpoint, Fortigate, Palo Alto, etc.).Scripting Python/Go/similarWillingness to participate in a rotating on-call schedule ...

Senior Network Architect - Low latency trading

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Linux, covering boot sequence, service management, file, process, network, and kernel management. Ability to deploy and operate monitoring frameworks such as Cacti, Zabbix, Nagios, Prometheus, or Tick stack. Strong firewall experience (Checkpoint, Fortigate, Palo Alto, etc.).Scripting Python/Go/similarWillingness to participate in a rotating on-call schedule ...

Senior Network Architect - HFT

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Linux, covering boot sequence, service management, file, process, network, and kernel management. Ability to deploy and operate monitoring frameworks such as Cacti, Zabbix, Nagios, Prometheus, or Tick stack. Strong firewall experience (Checkpoint, Fortigate, Palo Alto, etc.).Scripting Python/Go/similarWillingness to participate in a rotating on-call schedule ...

Software Engineer II - Backend (Ruby)

Location
Greater London, England, United Kingdom
grade systems (e.g. you know your way around logs, traces, metrics, feature flags, and can debug live traffic issues)* Familiarity with observability tooling (e.g. Prometheus, Grafana, Sentry and Lightstep)* A good understanding of Domain Driven Design* experience integrating with third-party APIs and services, particularly those with nuanced state transitions ...

Senior DevOps Engineer

Location
Ashford, England, United Kingdom
tooling, automation and infrastructure that power secure, high‐volume payment services used globally. You’ll work hands‐on with technologies like Kubernetes, ELK, Prometheus, Python/Go and modern CI/CD frameworks to shape how software is delivered at scale. Your impact will be visible: from improving observability … with DevOps/Platform Engineering within production environments Expert knowledge of operating, configuring and optimizing enterprise monitoring systems such as ELK, New Relic, Grafana, Prometheus etc. Strong knowledge of Kubernetes Strong programming/scripting skills in a modern language like Python, Go, or Rust Strong, hands‐on knowledge of DevOps ...

Cloud Infrastructure Engineer

Location
Greater London, England, United Kingdom
infrastructure — monitoring spend, eliminating waste, and rightsizing resources to balance performance and cost. Monitoring & Incident Management Monitor and manage platform activity using tools like Prometheus , Grafana , or AWS CloudWatch Respond quickly to alerts and incidents, independently resolving issues and ensuring service uptime. Conduct post‐incident reviews and help improve system … Strong experience with AWS services and containerised applications Strong experience operating operational data stores (Aurora MySQL, DynamoDB). Expertise in using monitoring tools(e.g. Prometheus, Grafana, CloudWatch) for real‐time platform performance insights. Strong understanding of network security and Cloudflare, VPC and networking fundamentals, with a clear grasp ...

Cloud Engineering & Architecture - Senior Platform Engineer AI - Vice President

Location
Greater London, England, United Kingdom
Temporal, or custom agentic loops) to coordinate multi-step diagnostic and remediation tasks. AIOps & Intelligent Observability: Ability to integrate traditional observability stacks (e.g., Datadog, Prometheus, OpenTelemetry) with AI/ML models to automate root-cause analysis, anomaly detection, and semantic log clustering. Self-healing Infrastructure Engineering: Experience designing closed-loop … Experience working in regulated industries is a plus. Preferred Qualifications Experience building self-service platforms for development teams. Familiarity with observability and monitoring tools (Prometheus, Grafana, Datadog, CloudWatch). Background in financial services or other highly regulated environments. About Goldman Sachs At Goldman Sachs, we commit our people, capital ...

Senior Data Engineer

Location
Altrincham, England, United Kingdom
quality checks as first‐class concerns. Ensure data quality and observability: Embed data testing and validation processes alongside monitoring and alerting systems such as Prometheus and Grafana to ensure reliability, performance, and operational insight across data services. Work with metadata and standards: Implement metadata capture and validation within data pipelines … checks throughout pipelines to ensure accuracy, completeness, consistency, and reliability. Practical experience implementing monitoring, observability and alerting for data platforms using tools such as Prometheus and Grafana. Strong understanding of data architecture patterns, including data lakes, data warehouses, and event‐driven architectures. Awareness of GDPR and data protection practices, including ...

Sr. Technical Account Manager - EMEA

Hiring Organisation
Wise
Location
London, UK
Employment Type
Full-time
financial infrastructure or another API-first technology environment. Have experience with SQL or database querying and observability or monitoring tools such as Kibana, Looker, Prometheus or similar. Have coding or scripting experience and are comfortable reading technical implementations to understand how systems behave. Have experience troubleshooting back-end service issues ...

Senior DevOps Engineer

Hiring Organisation
Granite Recruitment and Consulting
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
reliable services. You will get to work in a modern, cloud-based tech environment, and work with technologies such as Azure, Kubernetes, Terraform, Chef, Prometheus and Grafana. The team are passionate about working with cutting edge technology and adopt the latest DevOps/SRE methodologies. You will need the following …/CD tools and pipelines Azure DevOps or Argo CD It would be a bonus if you have experience with the following: Grafana Prometheus Chef Ansible Cloud platform security Benefits will include: 27 days holiday Private medical insurance Life assurance Flexible working A commitment to continue your professional development Career ...

Network Automation & OSS Designer

Location
Greater London, England, United Kingdom
pipelines. Architect AIOps capabilities including closed‐loop automation, anomaly detection, and predictive analytics for network operations. Integrate OSS observability with cloud‐native monitoring stacks (Prometheus, Grafana, OpenTelemetry, Elasticsearch). Lead design of intent‐based networking and policy‐driven automation frameworks. Collaborate with product managers, network engineers, platform teams, and DevOps … MANO, VNF/CNF lifecycle management). Understanding of AIOps platforms and closed‐loop automation design for network operations. Experience with observability tooling: OpenTelemetry, Prometheus, Grafana, Jaeger, Loki, ELK Stack. Knowledge of ML/AI model integration for anomaly detection, root‐cause analysis, and predictive network management. Strong grasp ...

DevOps & Infrastructure Engineer

Location
Gloucester, England, United Kingdom
solutions. Develop and maintain CI/CD pipelines, GitOps workflows and automated deployment approaches using tools such as ArgoCD. Implement and improve observability using Prometheus, Grafana, logging and alerting to support resilient platform operations. Use infrastructure-as-code and platform automation with Helm, Go and Terraform to deliver repeatable, assured …/CD and GitOps tooling experience, ideally including ArgoCD and automated deployment pipelines. Good understanding of observability, monitoring and alerting using tools such as Prometheus and Grafana, alongside security, networking, logging, secrets management and operational assurance. Able to learn new technologies quickly and help others adopt them safely and effectively. ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
integrations with third-party custodians. Build out Kubernetes and Docker deployments as we containerize more of the custody stack. Set up monitoring and alerting (Prometheus, Grafana, or equivalent) so we know about problems before customers do. Apply security controls and standards throughout - access control, key management, incident response. Provide … skills, with real production experience. A strong security focus - you've worked on systems where key management and access control are critical. Experience with Prometheus, Grafana, or an equivalent monitoring stack. Experience integrating with third-party custodians like BitGo or Fireblocks - this is a key differentiator for this role. What ...

Senior Software Engineer, Custody Services

Hiring Organisation
Robinhood Financial
Location
London, UK
Employment Type
Full-time
integrations with third-party custodians. Build out Kubernetes and Docker deployments as we containerize more of the custody stack. Set up monitoring and alerting (Prometheus, Grafana, or equivalent) so we know about problems before customers do. Apply security controls and standards throughout — access control, key management, incident response. Provide … skills, with real production experience. A strong security focus — you've worked on systems where key management and access control are critical. Experience with Prometheus, Grafana, or an equivalent monitoring stack. Bonus pointsExperience integrating with third-party custodians like BitGo or Fireblocks — this is a key differentiator for this role. ...