1 to 25 of 126 Permanent Datadog Jobs

Senior DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
fundamentals (DNS, routing, load balancing, VPNs, firewalls). Familiarity with monitoring and logging tools (e.g. Prometheus, Grafana, ELK/EFK stack, CloudWatch, Azure Monitor, Datadog, etc.). Good understanding of security best practices in cloud and Linux environments (IAM, least privilege, secrets management, patching). Experience presenting technical concepts ...

Cloud Platform Architect (68020)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Azure DevOps for infrastructure pipelines Policy‐as‐code frameworks: OPA/Rego, HashiCorp Sentinel, Azure Policy, AWS Config Rules Monitoring and observability: Prometheus, Grafana, Datadog, CloudWatch, or Dynatrace Networking fundamentals: VPC/VNet design, load balancers, DNS, CDN, and hybrid connectivity Soft skills & Competencies Strong leadership and mentoring ability — coaches ...

Cloud Platform Architect (68020)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Azure DevOps for infrastructure pipelines Policy‐as‐code frameworks: OPA/Rego, HashiCorp Sentinel, Azure Policy, AWS Config Rules Monitoring and observability: Prometheus, Grafana, Datadog, CloudWatch, or Dynatrace Networking fundamentals: VPC/VNet design, load balancers, DNS, CDN, and hybrid connectivity Soft Skills & Competencies Strong leadership and mentoring ability – coaches ...

Lead Software Engineer, Middleware Reliability Engineering

Hiring Organisation
Visa Technology and Operations LLC
Location
Foster City, California, United States
Employment Type
Permanent
Salary
USD Annual
systems, networking protocols, certificate management, secret management, system design, cloud platforms (AWS, • Azure, GCP), and containerization (Kubernetes, Docker • Proficiency with monitoring tools (Prometheus, Grafana, Datadog, etc.), logging systems (ELK stack, Splunk), and tracing tools (Jaeger, Zipkin). • Proficiency in infrastructure-as-code tools such as Terraform and Ansible. • Hands ...

Cloud Platform Architect (68020)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Azure DevOps for infrastructure pipelines Policy‐as‐code frameworks: OPA/Rego, HashiCorp Sentinel, Azure Policy, AWS Config Rules Monitoring and observability: Prometheus, Grafana, Datadog, CloudWatch, or Dynatrace Networking fundamentals: VPC/VNet design, load balancers, DNS, CDN, and hybrid connectivity SOFT SKILLS & COMPETENCIES Strong leadership and mentoring ability – coaches ...

Principal/Senior Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Welwyn, England, United Kingdom
chaos engineering experiments that validate systems and surface weaknesses before incidents. Build deep observability with monitoring, logging, and alerting frameworks such as Prometheus, Grafana, Datadog, and ELK. Provide technical leadership to a team of engineers, fostering collaboration, innovation, and continuous improvement. Partner across teams to align infrastructure with ...

DevOps / Cloud / Platform Engineer (All Levels) - UK Wide

Hiring Organisation
describe.me
Location
London, South East, England, United Kingdom
Employment Type
Full-Time
Salary
£50,000 - £130,000 per annum
production—deployment, scaling, networking, troubleshooting CI/CD platforms (GitHub Actions, GitLab CI, Jenkins, CircleCI, ArgoCD or equivalent) Observability stack experience (Prometheus, Grafana, Datadog, ELK, OpenTelemetry, New Relic) Scripting and automation in Bash, Python or Go Linux fundamentals, networking basics and cloud security concepts Familiarity with secret management (Vault ...

Lead SRE - Data Product Studio

Hiring Organisation
17918
Location
Glasgow, Lanarkshire, United Kingdom
Experience designing, implementing, and operating observability practices (e.g., white/black box monitoring, SLO alerting, telemetry collection) using tools such as Grafana, Dynatrace, Prometheus, Datadog, or Splunk Hands-on experience designing and operating CI/CD pipelines and release automation (e.g., Jenkins, GitLab CI, GitHub Actions, Argo CD) Proficiency ...

Staff Site Reliability Engineer

Hiring Organisation
Visa Technology and Operations LLC
Location
Austin, Texas, United States
Employment Type
Permanent
Salary
USD Annual
incident response). • Understanding of database technologies including SQL, NoSQL, and data storage patterns. • Experience with observability stacks (Prometheus, Grafana, OpenTelemetry, ELK/EFK, Datadog, or similar). • Proficient in automation using Bash, Python, or Ansible-like tools. • Working knowledge of software engineering practices (version control, testing, code reviews, design ...

Senior Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Bash automation GitHub and CI/CD pipelines AWS and cloud‐native infrastructure Kubernetes/Amazon EKS Grafana stack, Prometheus, Loki or Datadog JFrog Artifactory or artifact management Docker or container runtime experience AWS Batch, Step Functions, IAM and Karpenter Hybrid cloud platform migration or modernisation Secure platform design ...

Senior Site Reliability Engineer

Hiring Organisation
Spectrum It Recruitment Limited
Location
London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£65,000
cloud infrastructure Kubernetes and Docker Production support and incident management Python, Bash or Go scripting Monitoring and observability platforms such as Grafana, Prometheus, Datadog, Splunk or CloudWatch Networking fundamentals including DNS, TCP/IP and load balancing A passion for automation, continuous improvement and operational excellence Experience with Infrastructure ...

Senior Lead Site Reliability / DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab, Terraform, etc.) Experience with container and container orchestration (e.g., ECS, Kubernetes ...

Site Reliability Engineer - Comcast Technology Solutions

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
with infrastructure‐as‐code tools (e.g., Terraform, Ansible). Familiarity with containerization and orchestration tools (e.g., Docker, Kubernetes). Experience with monitoring tools (e.g., Datadog, Splunk). Experience with database performance monitoring and tuning (e.g., NoSQL, SQL). Experience with Kubernetes performance monitoring and tuning. Excellent problem‐solving skills ...

Senior Systems Engineer, Production

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
infrastructure at scale .Strong expertise with Terraform and infrastructure automation.Solid understanding of AWS services such as ECS/EKS, RDS, LambdaFamiliarity with observability platforms (Datadog, Prometheus, Grafana, etc.).Proficiency in containerization (Docker, Kubernetes).Experience with CI/CD systems such as Buildkite, GitHub Actions, etc.Familiarity with Linux systems administration , networking ...

Senior Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Experience owning, managing, and maintaining mission‐critical operational tooling. Desirable: Proven background in implementing and managing centralised logging solutions or similar platforms (e.g., Splunk, DataDog). Desirable: Familiarity with distributed tracing tools (e.g., Jaeger, Zipkin) and Application Performance Monitoring (APM) solutions. What we offer Competitive salary ...

Site Reliability Engineer (AWS)

Hiring Organisation
Spectrum IT Recruitment
Location
Birmingham, West Midlands, West Midlands (County), United Kingdom
Employment Type
Permanent
cloud infrastructure Kubernetes and Docker Production support and incident management Python, Bash or Go scripting Monitoring and observability platforms such as Grafana, Prometheus, Datadog, Splunk or CloudWatch Networking fundamentals including DNS, TCP/IP and load balancing A passion for automation, continuous improvement and operational excellence Experience with Infrastructure ...

Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment
Location
Basingstoke, Hampshire, United Kingdom
Employment Type
Permanent
reliability objectives (SLIs/SLOs). Reduce alert fatigue through continuous tuning and optimisation. Build and maintain dashboards using technologies such as: Grafana Prometheus Datadog Splunk AWS CloudWatch Reliability Engineering & Automation Automate repetitive operational tasks to minimise manual effort. Improve Mean Time to Detect (MTTD) and Mean Time to Resolve ...

Senior Engineer (Platform Engineering)

Hiring Organisation
17918
Location
London, United Kingdom
Secrets Operator, Vault, or similar secret sync patterns is beneficial Familiarity with Terraform, multi cloud environments (GCP and Azure), observability tooling (e.g. New Relic, Datadog, Prometheus, Grafana), or contract testing with Pact would be valuable ...

Senior Lead Site Reliability Engineer

Hiring Organisation
17918
Location
Glasgow, Lanarkshire, United Kingdom
proficiency and experience in observability such as white and black box monitoring, SLO alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, Elasticsearch, etc. Proficiency in continuous integration and continuous delivery tools (e.g., Jenkins, GitLab, Terraform, etc.) Experience with container and container orchestration (e.g., ECS, Kubernetes ...

SME - AzureMSSQL DBA, Microsoft Azure

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Database vendor certifications (Microsoft, Oracle, MySQL, PostgreSQL). Familiarity with CI/CD pipelines and DevOps practices. Experience with monitoring tools such as CloudWatch, Datadog, Grafana. Knowledge of containerization (Docker, Kubernetes) and exposure to data warehousing concepts and ETL processes. What We Offer Opportunity to lead and shape database strategy ...

Software Engineer, Infrastructure Reliability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Terraform. Proficiency in programming/scripting languages. Experience with containerization technologies and container orchestration platforms like Kubernetes. Experience with observability tools such as Datadog, Prometheus, Grafana, Splunk, and ELK stack. Experience with microservices architecture and service mesh technologies. Knowledge of security best practices in cloud environments. Strong understanding of distributed ...

Architect & Delivery Lead (68018)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
delivery models Financial modelling: FinOps, TCO analysis, business case development, and outcome‐based commercial models Tooling breadth: ITSM (ServiceNow), monitoring (Grafana/ELK/Datadog/Dynatrace), automation (Ansible/Terraform), and collaboration platforms Soft Skills & Competencies Exceptional leadership presence—inspires confidence at C‐level while earning respect from engineering ...

SVP of Engineering

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
engineering, function calling, agent frameworks) and graph/knowledge graph technologies. DevOps/SRE practices at scale: CI/CD, IaC (Terraform, Pulumi), observability (Datadog, Grafana), incident management. Leadership Qualities Builder mentality with hands-on orientation; executive presence and strong communication skills; collaborative and bias for action; comfort with ambiguity. ...

Staff Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
builders. Bonus Points Deep experience with Google Cloud Platform (GCP) services and tools. Expert-level knowledge of modern observability platforms (e.g., Prometheus, Grafana, Datadog, OpenTelemetry). Experience designing and building reliable systems capable of handling high throughput and low latency. Significant experience with Go and Terraform. Familiarity with working ...

AI Platform engineer

Hiring Organisation
Nextech Group Limited
Location
East London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
Infra: AWS (ECS, Lambda, SQS/SNS), Docker, Kubernetes, Terraform AI/ML tooling: LangChain/LlamaIndex, vLLM, Anthropic & OpenAI APIs, embedding models Observability: Datadog, Grafana, OpenTelemetry CI/CD: GitHub Actions, ArgoCD Requirements: 4+ years backend development experience, ideally with at least 1 year working with LLM/ ...