101 to 125 of 175 Remote/Hybrid Prometheus Jobs

Senior Backend Engineer | AI Platform

Location
Greater London, England, United Kingdom
organization. Tech Stack: Backend Python FastAPI Agent Development Kit (ADK) Datastores PostgreSQL BigQuery Firestore Infrastructure Google Cloud Platform (GCP) RabbitMQ Terraform Monitoring & Observability Grafana Prometheus Langfuse incident.io Sentry What to Expect from Our Hiring Process At Plum, we value a lot the time you devote to the hiring process, this ...

Site Reliability Engineer - NS London

Hiring Organisation
BAE SYSTEMS
Location
London, UK
Employment Type
Full-time
Mongo, Postgreso Know your way around Linux and Windows command lines, e.g. Bash and PowerShello Monitoring large systems using technologies such as Grafana, Prometheus, ELK, Splunko Experience of working in Agile teams, and the tooling that supports it, e.g. Atlassiano Diagnosing and troubleshooting application issues resulting in service outageso Troubleshooting ...

Senior Software Engineer - Space Reliability

Hiring Organisation
Spire Global
Location
Glasgow, UK
Employment Type
Full-time
Postgres, Redis, Elasticsearch, or S3Hands-on experience with Databricks or modern data Lakehouse toolingExperience building or operating monitoring and alerting systems such as Grafana, Prometheus, or NagiosInfrastructure as Code experience with tools such as Terraform or AnsibleML or AI applications in an operational or monitoring contextExperience with Python data visualization ...

Senior Software Engineer, Video Encoding

Location
United Kingdom
Experience with GPU‐accelerated encoding or hardware media pipelines Familiarity with Kubernetes, ECS, Nomad, or other orchestration platforms Experience with observability stacks such as Prometheus, Grafana, OpenTelemetry, ELK, or Datadog Experience building fault‐tolerant ingest or transcoding platforms operating across multiple regions Our Hybrid Work Approach Roku fosters an inclusive ...

Staff Platform Engineer

Location
Greater London, England, United Kingdom
Fargate clusters in AWS, creating common tooling to aid in development tasks, and running shared services such as Opensearch, Envoy, Vault and Prometheus to name a few. The team has also expanded its scope to simplify Data engineering in the organisation using the same techniques we used to ease creating ...

Platform Engineer

Location
Greater London, England, United Kingdom
Evaluation & Quality: Eval harnesses and golden datasets, LLM-as-judge and human-in-the-loop review, regression suites, and red-teaming Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, including GenAI tracing and token, latency, and cost telemetry Platform Security & Policy-as-Code: HashiCorp Vault, OPA/Conftest … supporting cloud or Kubernetes resources. Observability, Monitoring & Site Reliability (SRE) Instrument services and implement monitoring, logging, and alerting as code using standard tooling (Prometheus, Grafana, OpenTelemetry). Participate in the on‐call rotation, responding to incidents and helping restore service. Contribute to blameless post‐incident reviews and implement follow ...

Site Reliability Engineer III

Location
Belfast City District, Northern Ireland, United Kingdom
Manage cluster lifecycles, data replication, RBAC, and workload placement. Observability & Monitoring Fabric: Design, scale, and maintain our observability backbone using tools like OpenTelemetry, Splunk, Prometheus, and Grafana. Establish and continuously improve metrics, logs, alerting strategies, SLIs, and SLOs to enable fast issue detection. Incident Response & Operations: Engage with urgency … with an eagerness to learn independently and collaboratively. Preferred Qualifications/Desirable Observability Stack: Hands-on experience with telemetry tools such as OpenTelemetry, Splunk, Prometheus, and Grafana. Agile Integration: Comfort working within Agile frameworks and collaborative software development lifecycles. Certifications: GCP Professional Cloud Architect, Certified Kubernetes Administrator (CKA), or Certified ...

Cloud Operations Engineer (remote - London)

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Experienced and Certified in cloud computing with AWSExperience in a public cloud such as AWSKnowledge of monitoring and alerting technologies such as Grafana, Prometheus,Expereince of working with of Docker & Kubernetes and Container technology in productionWindows and Linux Operating System Management TechniquesSolid understanding of the OSI ModelExperience in database technology ...

Site Reliability Engineer

Location
City Of London, England, United Kingdom
speed and reducing deployment risk. Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments ...

Senior Site Reliability Engineer

Hiring Organisation
CISCO Systems
Location
London, UK
Employment Type
Full-time
speed and reducing deployment risk. Adaptable & Problem-Solver: Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality: Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments ...

Non-Functional Test Specialist

Hiring Organisation
Euroclear
Location
United Kingdom, UK
Employment Type
Full-time
experience in planning, preparing & executing Resilience/Disaster Recovery/Continuity/Failover testing/PerformanceExposure to tools supporting resilience & operational observability (Splunk, Prometheus, Kafka, Chaos tooling etc)Understanding of high-availability architectures, infrastructure redundancy & backup/restore strategiesAre familiar with working within large scale & complex Technology Implementation or changeExperience ...

Neo4j Platform Consultant

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
scaling)Security (RBAC, authentication, data protection)Preferred SkillsExperience with Graph RAG workloads and traversal optimizationDevOps tools (Docker, Kubernetes, CI/CD pipelines)Monitoring tools (Prometheus, Grafana, etc.)Experience with other graph databases (Neptune, TigerGraph)Job ExpectationsEnsure stable, secure, and high-performing Neo4j platform operationsEnable efficient graph query execution ...

Engineer II, Site Reliability (Hybrid, London)

Location
Greater London, England, United Kingdom
drive to make things better Bias towards small development projects and the occasional larger projects Have experience with modern monitoring and telemetry stacks (ELK, Prometheus, Grafana, Zabbix) Gather and analyze metrics from both operating systems and applications to assist in performance tuning and fault finding Ability to lead incident analysis ...

Software Engineer (ML Projects)

Location
Greater London, England, United Kingdom
cloud‐native TeamCity for CI/CD (lots of teams are releasing code 15-20 times per day!) Terraform Prometheus and Grafana If you have built and deployed complex Python applications or have hands‐on experience with generative AI and LLMs, we would be especially keen to talk. ...

Senior Software Engineer - Live & VOD Video Infrastructure

Hiring Organisation
Roku
Location
Cambridge, Cambridgeshire, UK
Employment Type
Full-time
similar technologiesExperience with GPU-accelerated encoding or hardware media pipelinesFamiliarity with Kubernetes, ECS, Nomad, or other orchestration platformsExperience with observability stacks such as Prometheus, Grafana, OpenTelemetry, ELK, or DatadogExperience building fault-tolerant ingest or transcoding platforms operating across multiple regions#LI-JC5What's Roku's approach to hybrid working? Roku fosters ...

Database Reliability Engineer

Location
Manchester, England, United Kingdom
multi-cloud—while ensuring rigorous data integrity and mobility A Security & Observability Mindset: You believe security is paramount. You focus on building deep observability (Prometheus/Grafana/OpenTelemetry/Humio) and automated guardrails so the fleet is secure by design without requiring manual intervention Engineering via Code: While ...

Test Environment Manager (10105)

Location
Greater London, England, United Kingdom
Improvement Monitor environment availability, health, performance and utilisation. Develop appropriate metrics and dashboards for environment reporting. Work with monitoring and logging technologies such as Prometheus, Grafana and Splunk. Identify opportunities to improve environment reliability, automation, scalability and cost efficiency. Drive continuous improvement across Test Environment Management processes. Essential Experience 5+ … Operations teams. Highly Desirable Experience Azure/Azure DevOps Jenkins/GitLab Terraform or other IaC technologies HP NonStop infrastructure Java-based environments Prometheus/Grafana/Splunk UFT/Selenium/Cucumber Linux shell scripting Large-scale financial services or similarly complex regulated environments *Rates depend on experience ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
Job Title: Java Kafka EngineerLocation: NorthamptonAbout the Job you are considering:Join the Financial Services Business Unit within Capgemini’s Cloud & Custom Applications (C&CA) practice, where we Consult with Purpose, Continuously Evolve, and Architect ...

Software Engineering Team Lead – Full Time – Cardiff

Location
Cardiff, Wales, United Kingdom
practices Desirable: Financial Services/FinTech experience React and TypeScript CI/CD tools such as Jenkins Monitoring and observability tools including Grafana or Prometheus Interest in AI-enabled software development Benefits: Salary up to £75,000 Performance-related bonus Hybrid working Private healthcare Enhanced annual leave Employee wellbeing programme ...

Senior Architect Private Cloud

Hiring Organisation
Randstad Technologies Recruitment
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£500 - £550/day Inside IR35 via Umbrella
Skills: Familiarity with Service Mesh (e.g., Istio), API Gateways, Policy as Code (e.g., OPA), developer portals (Platform-as-a-Product), and Observability stacks (OpenTelemetry, Prometheus, ELK). Randstad Technologies is acting as an Employment Business in relation to this vacancy. ...

Software Engineer (EMS)

Location
Greater London, England, United Kingdom
Azure) and tools for multi‐region deployments. A working knowledge of databases and caches (PostgreSQL, SQLite, Redis, Zookeeper). Experience with observability tools (e.g. Prometheus, Grafana, ELK stack). Soft Skills: Strong analytical and problem‐solving abilities. Excellent communication and collaboration skills, with a track record of working in cross ...

Senior .NET Backend Developer

Location
York and North Yorkshire, England, United Kingdom
handling sensitive or clinically important data. PostgreSQL, Redis, Elasticsearch or other data and caching technologies. GraphQL, including schema design and gateway patterns. Grafana, OpenTelemetry, Prometheus or equivalent observability tooling. CI/CD, production services, microservices and message-driven systems such as RabbitMQ. How we work We value engineers ...

Staff Python Engineer (ML)

Location
City Of London, England, United Kingdom
clear, and easy to test Developing observability for new and existing ML applications and GenAI/LLM integrations , making use of the Grafana Stack (Prometheus, Loki, Tempo) Develop integrations and services that communicate with Google Services. Working closely with Data Scientists and ML Engineers throughout the lifecycle of productionising their ...

Senior Software Engineer (Payments)

Hiring Organisation
ebury
Location
London, UK
Employment Type
Full-time
should be excited to work with: Languages/Frameworks: Python, Django, FastAPIData/Messaging: PostgreSQL, KafkaInfrastructure/DevOps: AWS, Kubernetes, Terraform, JenkinsObservability: Prometheus, KibanaWhat We OfferCompetitive salary and benefits packageDiscretionary performance-based bonusA truly dynamic, global environment with massive scale and exposureContinued personal development through structured training and certificationsEqual Opportunity ...

Database Reliability Engineer

Hiring Organisation
Starling Bank
Location
London, UK
Employment Type
Full-time
multi-cloud—while ensuring rigorous data integrity and mobilityA Security & Observability Mindset: You believe security is paramount. You focus on building deep observability (Prometheus/Grafana/OpenTelemetry/Humio) and automated guardrails so the fleet is secure by design without requiring manual interventionEngineering via Code: While ...