1,801 to 1,825 of 5,503 Permanent Observability Jobs

Lead Site Reliability Engineer (Dynatrace)

Hiring Organisation
SF Partners
Location
South West England, United Kingdom
Employment Type
Full-Time
Salary
£80,000 - £100,000 per annum
looking for an experienced Site Reliability Engineer/Observability Engineer with deep Dynatrace expertise to join a major technology and platform engineering programme. This is not a role for someone who has simply used Dynatrace dashboards. We're looking for an engineer who has been involved in the implementation, configuration … technical SME within complex production environments. What we're looking for Strong hands-on Dynatrace implementation and administration experience Experience designing and implementing observability/monitoring solutions end-to-end Strong SRE and production engineering background Experience configuring instrumentation, metrics, alerting and monitoring Understanding of technologies such as OneAgent, ActiveGate ...

Fullstack Engineer

Location
Bracknell, England, United Kingdom
Design and maintain scalable Go-based microservices Build and support REST and gRPC APIs Develop integrations and event-driven solutions across distributed systems Improve observability, reliability, scalability, and security Deploy and support applications in AWS Work with Docker, Kubernetes, and CI/CD pipelines Participate in production support … Nice to have: AWS experience Nice to have: Kubernetes and container orchestration Nice to have: Event-driven architectures and messaging platforms Nice to have: Observability, monitoring, and distributed tracing Nice to have: Experience working in a SaaS product organisation Core Competencies Demonstrates expertise in building modern web applications using React ...

Senior SRE Engineer

Hiring Organisation
SF Partners
Location
Birmingham, West Midlands (County), United Kingdom
Employment Type
Permanent
Salary
£100000 - £110000/annum great training and progression opp
largest and most complex enterprise platforms. Working within a highly skilled, multi-disciplinary engineering team, this Senior SRE Engineer will take technical ownership of observability and reliability across large-scale AWS environments, helping engineering teams improve platform performance, resilience and availability. A key focus of the role will … enterprise environments - End-to-end Dynatrace implementation experience - Experience designing and deploying Dynatrace across complex cloud environments - Strong understanding of APM, infrastructure monitoring and observability - Dynatrace configuration, dashboards, alerting and performance monitoring - Experience integrating Dynatrace into AWS and wider engineering/tooling ecosystems - Strong understanding of SLIs, SLOs, availability, reliability ...

Site Reliability Engineer with Python

Hiring Organisation
BC Forward
Location
Charlotte, North Carolina, United States
Employment Type
Permanent
Salary
USD Hourly
Site Reliability Engineer with Python to join our dynamic team. The ideal candidate will have strong experience in Python development, Linux administration, infrastructure automation, observability, and cloud-native platforms and a proven ability to improve reliability, enhance observability, and drive operational efficiency at scale. Responsibilities: Own reliability and operational health … core platform services and components. Design, implement, and maintain automation for infrastructure and operational workflows. Improve platform and application observability using modern monitoring and logging tools. Lead incident response, troubleshoot production issues, and drive post-incident remediation. Partner with core engineering teams on platform modernization and cloud migration initiatives. Harden ...

SRE Engineer

Location
Greater London, England, United Kingdom
automate the deployment of our software Automate the provisioning and management of our infrastructure using Infrastructure as Code (IaC) tools Define, implement, and maintain observability solutions for our applications to ensure we can proactively detect system degradation, easily understand system state, and quickly diagnose issues Diagnose and resolve production issues … must have and one should be good at coding in terraform CICD Tools hands on : Jenkins , GitHub , GitHub Actions, Cloud Deployment pipelines Observability Tools - Splunk//Graphana/Datadog and Distributed Tracing, ELF, & Dynatrace Problem-Solving : Proven ability to troubleshoot complex issues in distributed systems and debug problems effectively. ...

Machine Learning Operations Engineer

Location
Greater London, England, United Kingdom
build and operate the platform capabilities that take machine-learning models from experimentation into reliable production services You’ll own the automation, deployment, observability and operational controls around the ML lifecycle, working closely with research engineers, software engineers, platform teams and product teams This is not a research role. … model metadata and reproducibility across research and production Build reusable tooling and platform capabilities that support multiple models and engineering teams Model serving and observability Deploy and operate batch and online inference services in containerised cloud environments Define and meet availability, latency, throughput and recovery objectives for ML services Monitor ...

Software Engineer, Real-Time

Location
Nottingham, England, United Kingdom
that show how reliably and quickly market data reaches customers. You will work with experienced engineers to develop production software, measure performance and enable observability of a globally distributed platform. Prior market-data or observability experience is not required; we are looking for strong engineering fundamentals, curiosity and a willingness … hands‐on development role for an engineer at an early stage of their career who wants to build experience in mission-critical distributed systems, observability, cloud-native engineering and real‐time financial technology. What You’ll Be Doing Build and improve software components that aggregate, correlate and present data from ...

Director of Site Reliability Engineering

Location
Greater London, England, United Kingdom
robust incident management frameworks and lead major incident response activities for critical systems Implement blameless postmortems and deliver systemic improvements across production environments Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets … operational tooling to reduce manual processes Requirements Strong background in Site Reliability Engineering, DevOps, or platform operations in complex, distributed environments Expertise in observability platforms, troubleshooting distributed systems, and telemetry‐driven insights Hands‐on experience with automation, Infrastructure as Code (Terraform or CloudFormation), and CI/CD practices Deep understanding ...

Software Engineer, Real-Time

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
that show how reliably and quickly market data reaches customers. You will work with experienced engineers to develop production software, measure performance and enable observability of a globally distributed platform. Prior market-data or observability experience is not required; we are looking for strong engineering fundamentals, curiosity and a willingness … hands-on development role for an engineer at an early stage of their career who wants to build experience in mission-critical distributed systems, observability, cloud-native engineering and real-time financial technology. WHAT YOU'LL BE DOINGBuild and improve software components that aggregate, correlate and present data from over ...

Senior DevOps Engineer

Location
East Midlands, England, United Kingdom
/CD pipelines and deployment orchestration. Support Kubernetes and OpenShift platform troubleshooting and optimisation. Deliver secure and compliant infrastructure solutions. Implement monitoring, logging, and observability tooling across environments. Collaborate with engineering, architecture, and delivery teams to improve deployment efficiency and platform reliability. Champion automation-first approaches to infrastructure and application … designing and maintaining enterprise-scale CI/CD pipelines . Strong understanding of cloud security and secure delivery practices. Experience implementing monitoring, logging, and observability solutions. Ability to define technical standards, governance, and reusable deployment frameworks. Experience working within large-scale enterprise transformation programmes. Desirable Skills Experience within highly regulated ...

Cloud SRE — IaC, Kubernetes & CI/CD

Location
England, United Kingdom
will manage Kubernetes clusters (EKS/AKS), build CI/CD pipelines with GitHub Actions, and implement monitoring with Prometheus and Grafana to ensure observability, security and cost-efficiency across platforms. #J-18808-Ljbffr ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
Hertfordshire, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
Surrey, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
Merseyside, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
Lincolnshire, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
Cornwall, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
Anglesey, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
Monmouthshire, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
South yorkshire, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
South ayrshire, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
Antrim and newtownabbey, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
Redcar and cleveland, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Remote Senior AI Software Engineer

Hiring Organisation
Aveni
Location
Stockton-on-tees, United Kingdom
integrations that bring those models to life for advisers, analysts, and banking teams. You'll own features end-to-end, covering selection, implementation, deployment, observability, and iteration across backend, frontend, and cloud infrastructure. Key Responsibilities Build production-ready, AI-powered products integrating LLM capabilities, RAG, and agentic workflows into real … world processes. Select models, orchestrate prompts, design systems, and implement AI tooling (e.g., LangGraph, Langfuse). Design evaluation, monitoring, and observability to ensure reliability and readiness for production. Develop scalable, event-driven microservices and APIs using Node.js and TypeScript. Build modern, responsive React frontends to make AI useful in customer ...

Senior Platform Engineer: AI-Ready Infra & Security

Location
Slough, England, United Kingdom
scalable platform features, automate operations, and build self-service workflows. The role emphasizes strong Python, Linux, Terraform/Ansible, Docker and Kubernetes proficiency, plus observability with Prometheus, Grafana and OpenTelemetry. #J-18808-Ljbffr ...