151 to 175 of 237 Observability Jobs in the South East

Data Scientist & Engineer

Hiring Organisation
Norton Rose Fulbright LLP
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
contracts and reusable features for entities such as clients, organisations, people, matters, sectors, jurisdictions, opportunities and legal topics. Implement data-quality checks, lineage, provenance, observability, monitoring and change management within Fabric and Databricks workflows so data assets remain trusted over time. Conduct exploratory data analysis to surface patterns, gaps, anomalies ...

Senior Forward Deployment Engineer

Location
Slough, England, United Kingdom
analysis, upgrading Java and NPM runtimes, modernizing Spring and legacy middleware applications, improving CI/CD pipelines, containerizing applications, automating deployments, and introducing standard observability and resilience patterns. The Expert FDE is expected to lead complex engagements, work directly with development and client stakeholders, define the technical remediation approach, implement … testing, release, resilience, and legacy technology challenges with development teams. Assess application code, dependencies, runtime environment, test coverage, deployment architecture, CI/CD pipelines, observability, and operational risks. Write, debug, review, and enhance production-quality code and configuration throughout engagements. Define and implement practical modernization and remediation plans with clear ...

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
reliability, and availability Implement self-service tooling to empower development teams Drive DevOps best practices across the digital product lifecycle Develop and enhance monitoring, observability, and incident response processes Support global engineering teams delivering high-traffic platforms Key Requirements Proven experience supporting digital product delivery in a DevOps or platform … with Infrastructure as Code (Terraform, Ansible, Puppet or similar) Hands-on experience with Kubernetes, Docker, and cloud platforms (AWS preferred) Experience with monitoring/observability tools (Prometheus, Grafana, ELK, APM tools) Solid understanding of system performance, scalability, and resilience Strong collaboration and communication skills within cross-functional product teams Desirable ...

Full Stack Engineer - Must be DV Cleared

Hiring Organisation
VIQU IT Recruitment
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
platform engineers so models, pipelines and infrastructure land as one product rather than three. Improving the day to day: CI/CD, observability, security hardening, documentation. Sitting in front of customers and partners and turning what they need into something buildable. What You Will Need Python. Serious commercial experience building … secrets management. Putting AI or ML models into production, LLM features included. Infrastructure as code with Terraform, Ansible or similar, plus configuration management. Observability and incident response: metrics, logging, tracing. Message queues, event-driven architectures or data pipelines. Also Welcome Kubernetes. Deploying, running or debugging workloads, including K3s, RKE2 ...

Senior Site Reliability Engineer

Location
Reading, England, United Kingdom
inference infrastructure (GPU-backed endpoints, autoscaling, latency and cost tradeoffs), with SLOs, on-call, and incident response that cover models, not just services Observability(includes ML models) - drift and performance monitoring for ML, plus LLM-specific tracing, evals, and guardrails, wired into the same metrics and logging stacks … model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights & Biases, or equivalents) LLMOps in production - inference serving, prompt/version management, and LLM observability (tracing, evals, drift, guardrails, cost per request) Governing ML/LLM workloads as platform capabilities: data-residency and PII controls, and audit trails Any other ...

Monitoring & Observability Engineer

Location
Reading, England, United Kingdom
million Series C funding round – the largest fundraise ever completed by a quantum computing company in Europe. The Purpose As a Monitoring and Observability Engineer, you'll help keep OQC's live quantum computing systems running at their best. By improving system visibility, developing intelligent monitoring solutions and driving operational … Live Services team, you'll monitor the health and performance of our live cryogenic systems, respond to operational incidents and continuously improve our observability capabilities. You'll collaborate across engineering, operations, software and reliability teams to develop dashboards, refine alerting strategies and automate operational responses that improve reliability and reduce ...

Logs Specialist

Location
Maidenhead, England, United Kingdom
demonstrate the unique value of the Dynatrace GrailTM data lakehouse, helping customers transition from high-cost, fragmented logging silos to a unified, AI-powered observability platform.**Core Responsibilities****Domain Expertise:*** Act as the domain "subject matter expert" (SME) for Logs, staying ahead of industry trends like OpenTelemetry (OTel), log pipelines … management impacts MTTR and operational overhead.* Education: Bachelor's degree in - Computer science, Engineering, or equivalent practical experience.* Hands‐on exposure to modern observability pipelines, including Cribl solutions or OpenTelemetry logging specifications.* Familiarity with the Cribl ecosystem or OpenTelemetry logging specs.* Experience with scripting (Python, Go, or Bash ...

Cloud FinOps Analyst: Cost Optimisation & Governance

Location
Tonbridge, England, United Kingdom
financial accountability, cost transparency, and optimisation across cloud and data platforms, with a focus on Azure and Snowflake. The role involves leading governance, cost observability, and collaboration with Engineering, Data, Cloud Operations and Finance to enable a cost-efficient, data-driven cloud ecosystem in a hybrid office setup. #J ...

Senior Cloud Infrastructure Engineer (GCP)

Location
Southampton, England, United Kingdom
Infrastructure Engineer to design, build and run the cloud foundations underpinning our banking platform. You will own scalable cloud infrastructure across regions, emphasizing reliability, observability and security. You will lead complex projects, collaborate across teams, and help shape future capabilities with automation, policy-driven governance, and modern tooling ...

Field-Ready AI Engineer: Deploy & Scale Agentic AI

Location
Crawley, England, United Kingdom
production across diverse platforms. The role focuses on building production architectures, APIs, CI/CD pipelines and security practices, with emphasis on reliability, observability and scalable performance in scientific and enterprise contexts. #J-18808-Ljbffr ...

Data Platform Engineering Leader (Snowflake + Azure)

Location
Bexhill-on-Sea, England, United Kingdom
ways of working and Hastings strategy. You will drive a Centre of Excellence to support data teams, ensure data quality, security, lineage and observability, and influence decisions beyond technology for the company’s growth. #J-18808-Ljbffr ...

AI Platform Engineer — Remote-Eligible Platform Automation

Location
Reading, England, United Kingdom
engineers and security across the business to deliver scalable, secure platform capabilities. This hands-on engineering role covers platform delivery, AI-driven operations, resilience, observability and governance, with a focus on measurable outcomes and value across #J-18808-Ljbffr ...

Lead, Production Reliability & Ops

Location
Portsmouth, England, United Kingdom
Lead to own production and build scalable reliability. You’ll lead live customer-facing systems, shaping processes, incidents, and a team culture focused on observability and uptime. This hands-on role requires strong SRE/DevOps experience, proven leadership, and the ability to drive blameless postmortems and reliability roadmaps across ...

Fractional Lead Graph Engineer

Hiring Organisation
Elliot Marsh Head Hunting Partners
Location
London, South East England, United Kingdom
Employment Type
Part-Time
Salary
Salary negotiable
cases - Ensure graph data can be supplied efficiently and reliably to machine learning and inference services - Establish appropriate engineering standards for data integrity, resilience, observability and documentation - Work collaboratively with backend, AI and infrastructure specialists, transferring knowledge into the wider engineering team Fractional Lead Graph Engineer – You: - Significant commercial experience ...

Lead AI Architect - Agentic AI, Strategy & Adoption - Outside

Hiring Organisation
Sanderson
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£650.00 - £750.00 per day
enterprise-wide adoption. Key Responsibilities Define enterprise AI target architectures and transition roadmaps. Develop scalable AI and agentic architecture patterns covering orchestration, integration, security, observability and lifecycle management. Working with AI Governance, evolve the Agent taxonomy and tiering based on purpose, autonomy, impact, data sensitivity, reach and criticality. Including patterns ...

Software Engineer - Data & AI

Location
Woking, England, United Kingdom
source systems and downstream consumers at every integration boundary. Own the infrastructure your services run on: infrastructure-as-code, CI/CD, containerisation, and observability and deployment. Contribute to AI evaluation and quality: build and run automated evals and regression tests and use observability tooling to track latency, token usage ...

Advanced Engineer, Investment Technology

Location
Henley-on-Thames, England, United Kingdom
subject matter experts. Their primary focus is on execution within defined parameters, applying modern engineering practices including CI/CD, automated testing, code quality, observability and data quality controls. They may also be accountable for regular reporting or process administration within the squad. Working in partnership with more experienced staff … business requirements into practical, well-engineered technical solutions. Build production-ready solutions using modern engineering practices, including CI/CD, automated testing, code quality, observability and data quality controls. Support data quality, monitoring, reporting and process administration activities to help ensure accurate and consistent outcomes. Identify and resolve technical problems ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£67,547 - £83,778 per annum
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Fullstack Engineer

Location
Bracknell, England, United Kingdom
Design and maintain scalable Go-based microservices Build and support REST and gRPC APIs Develop integrations and event-driven solutions across distributed systems Improve observability, reliability, scalability, and security Cloud & DevOps Deploy and support applications in AWS Work with Docker, Kubernetes, and CI/CD pipelines Participate in production support … communication skills Nice to Have gRPC and Protocol Buffers MongoDB experience AWS experience Kubernetes and container orchestration Event-driven architectures and messaging platforms Observability, monitoring, and distributed tracing Experience working in a SaaS product organisation #UKJobs #SoftwareEngineering #Golang #TechjobsUK #LI-AD1 Flexera is proud to be an equal opportunity employer. ...

DevOps Engineer

Location
Reigate, England, United Kingdom
infrastructure. Automate environment provisioning across development and production. Manage backend state, pipelines, and state-change detection integrations. Platform Engineering & SRE Own and improve reliability, observability, and performance of the platform. Implement SLOs, alerting, dashboards, and auto remediation where possible. Troubleshoot cluster level, networking, and workload deployment issues. Lead root cause … endpoints, Certificate/Secret management etc Strong debugging and operational experience (SRE mindset). Solid experience of DevSecOps architecture, processes & tooling Solid understanding of Observability Process & Tooling Logging, metrics, traces, dashboards Other highly desirable, but not essential skills are: Experience with: GitOps - ArgoCD or GitOps workflows Zero downtime deployments (blue ...

SRE

Location
Hove, England, United Kingdom
No. of Positions: 1 We are seeking an experienced Site Reliability Engineer (SRE) to drive the modernization of IT operations through the implementation of observability practices, automation, and reliability engineering principles. The role requires a strategic thinker with strong hands‐on expertise who can enhance system reliability, scalability, and operational … practices, automate operational workflows, and establish robust monitoring and incident management frameworks. Key Responsibilities Collaborate with engineering teams to modernize IT operations by improving observability, automation, and operational efficiency. Design and implement observability platforms to effectively monitor system health, performance, and reliability. Develop strategies for AI-driven alerting and proactive ...

Production Reliability Lead

Location
Slough, England, United Kingdom
change coordination. In this hands-on leadership role, you’ll build a scalable operating model, foster strong runbooks, and drive continuous improvement across observability, MTTR, and incident prevention #J-18808-Ljbffr ...

Senior DevOps Engineer

Hiring Organisation
Halian Technology Limited
Location
Reading, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£95,000
Drive platform improvements and DevOps best practices. Design and implement self-service infrastructure and tooling. Deliver scalable, secure, and highly available systems. Enhance monitoring, observability, and operational performance. Support engineering teams with technical expertise and guidance. Skills & Experience Experience designing and implementing CI/CD pipelines and software delivery processes. … Infrastructure as Code experience using tools such as Terraform or Ansible. Experience with monitoring and observability tools. Strong knowledge of Docker, Kubernetes, AWS, and cloud technologies. Excellent communication skills and ability to collaborate across teams. A passion for automation, platform engineering, and continuous improvement. This is a full-time, permanent ...