1 to 25 of 47 Observability Jobs in Birmingham

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham Birmingham · United Kingdom · Vice President

Location
Birmingham, England, United Kingdom
ingress controllers. Advanced experience with major cloud providers (AWS, GCP, or Azure), specifically building and operating highly resilient cloud‐native architectures. Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch) Experience with automated testing and SDLC concepts, developing applications ...

The Core Engineering - Software Engineer - Associate - Birmingham

Location
Birmingham, England, United Kingdom
with automated testing and SDLC concepts, developing applications in a Linux environment, and sound knowledge of algorithms, data structures and software design. Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, ELK, or OpenTelemetry). Experience in some of the following is desired: SRE experience, relational ...

Senior Applied AI Engineer (Manager) TC

Location
Birmingham, England, United Kingdom
meet these demands. Role purpose Lead technical delivery across solution threads: set technical direction, mentor engineers, and ensure systems are production‐ready (reliability, observability, security, runbooks). Continue your development through the Applied AI Engineering Academy focused on advanced patterns and engineering leadership. EY Grade: Manager (UK) What ...

AI Native Software Engineer – Agentic & Applied AI

Location
Birmingham, England, United Kingdom
Building orchestration, tool invocation, routing and memory capabilities Developing evaluation frameworks and measuring accuracy, latency, cost and safety Implementing LLMOps practices including prompt versioning, observability and production monitoring Building APIs, backend services and full-stack applications that connect software systems with AI and agentic backends Designing and deploying cloud-native ...

Backend Developer - AWS/Typescript/Serverless Architecture/AWS Lambda

Location
Birmingham, England, United Kingdom
containerisation and Docker. Experience with event-driven architecture and messaging systems. Familiarity with security best practices for cloud-based applications. Experience with application monitoring, observability, logging and performance optimisation. Experience working with large-scale or public-facing digital services. Experience working within highly regulated or security-conscious environments. By joining ...

Senior Specialist Engineer (Specialist Site Reliability Engineer SRE)

Hiring Organisation
UK Health Security Agency
Location
Birmingham, Chilton, Leeds, Liverpool, London, Porton, E14 4PU, United Kingdom
Salary
£41983.00 to £52113.00
scalable, and perform optimally in production environments. The role will monitor and manage these aspects while taking responsibility for multiple cloud infrastructure services. Observability of systems will be key to prioritising the operational service improvements and performance improvements to meet and exceed SLOs (Service Level Objectives). Main duties … identify bottlenecks with an engineering mindset Ensure systems can handle current and future workloads through automation and capacity planning Continuously improve services through observability, and identify ways to improve observability practices Follow SRE principles. Guide and educate stakeholders to adopt implemented principles Provide technical documentation for engineers. Providing training, where ...

Senior DevOps Engineer

Hiring Organisation
Inspire People
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£50,000
technical leadership and subject matter expertise across Microsoft Azure * Build, maintain and optimise Infrastructure as Code solutions and CI/CD pipelines * Improve monitoring, observability and operational performance across digital services * Ensure the security, stability and availability of production and non-production environments Essential skills for the Senior DevOps Engineer ...

Internal Audit - Data Engineering - Associate - Birmingham

Location
Birmingham, England, United Kingdom
Snowflake and relational databases including DB2 and PostgreSQL, applying sound principles for schema design, performance and data movement. Implement data-quality checks, controls, lineage, observability, documentation and automated tests to improve trust and supportability. Contribute to AI-ready data products and semantic layers that provide consistent business meaning and governed ...

Senior Engineering Manager - SRE

Location
Birmingham, England, United Kingdom
systems, and how fast we can act on what we see. As Senior Manager, Site Reliability Engineering, you'll own the monitoring and observability roadmap across our full product portfolio — spanning Health, Legal, Education, Workforce Management, and other sectors — and lead a team of 15-20 SRE engineers to deliver … dashboard, and moving us toward a more autonomous operating model. Mandate: To lead OneAdvanced's Site Reliability Engineering function — building the monitoring and observability roadmap that gives every product across Health, Legal, Education, Workforce Management, and other sectors real visibility into its own health, and increasingly uses AI agents ...

Digital Engineer Developer - Agentic AI Engineer

Hiring Organisation
Hackajob Ltd
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent
agent patterns Understanding of integration patterns and enterprise systems Nice-to-Have Skills Experience with tools like n8n, LangGraph, or similar Knowledge of AI observability and testing frameworks Exposure to CI/CD and DevOps practices Experience in ITSM automation or managed services Experience Profile 38 years in software development ...

Remote Senior AWS Infra & DevOps Engineer

Location
Birmingham, England, United Kingdom
infrastructure behind a real-time health-tech platform. You will design and operate high-availability systems, implement IaC, and own observability and incident response. The role emphasizes security, automation, and scalable delivery. You will collaborate with engineering to optimize performance, cost, and resilience, with a path toward expanding ...

Senior Backend Engineer

Hiring Organisation
Inspire People
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
part of a multidisciplinary agile team, you will design, build and run platform services that underpin critical digital products, helping development teams improve observability, monitoring, CI/CD and service resilience. As a Senior Backend Engineer (Site Reliability), you will: * Design, build and operate reliable, secure and scalable cloud platform … services supporting critical digital products. * Develop and maintain platform tooling, automation, observability, monitoring and CI/CD capabilities. * Build software solutions using Python and modern engineering practices. * Write clean, maintainable code and infrastructure-as-code solutions to support service delivery. * Embed Site Reliability Engineering principles including SLIs, SLOs, error budgets ...

The Core Engineering - Site Reliability Engineering - Associate - Birmingham

Location
Birmingham, England, United Kingdom
operational resilience. Practice sustainable incident management through clear escalation, effective remediation, and a blameless postmortem culture. Identify and implement improvements to system behavior, controls, observability, and monitoring tools. Define and maintain service level indicators (SLIs), service level objectives (SLOs), and error budgets to quantify and manage service reliability. Engineer automation … development lifecycle concepts, and developing applications in a Linux environment. Strong understanding of algorithms, data structures, software design, and distributed systems fundamentals. Experience with observability platforms, including distributed tracing, logging, metrics, and tools such as Prometheus, Grafana, ELK, or OpenTelemetry. Experience with site reliability engineering practices, relational databases, Hadoop ...

The Core Engineering - Site Reliability Engineering - Associate - Birmingham

Location
Birmingham, England, United Kingdom
operational resilience. Practice sustainable incident management through clear escalation, effective remediation, and a blameless postmortem culture. Identify and implement improvements to system behavior, controls, observability, and monitoring tools. Define and maintain service level indicators (SLIs), service level objectives (SLOs), and error budgets to quantify and manage service reliability. Engineer automation … development lifecycle concepts, and developing applications in a Linux environment. Strong understanding of algorithms, data structures, software design, and distributed systems fundamentals. Experience with observability platforms, including distributed tracing, logging, metrics, and tools such as Prometheus, Grafana, ELK, or OpenTelemetry. Experience with site reliability engineering practices, relational databases, Hadoop ...

SRE Associate: Build Reliable Cloud Platforms

Location
Birmingham, England, United Kingdom
seeking a Site Reliability Engineer to join Core Engineering. You will help build, run and maintain high-performing, distributed systems, focusing on reliability, observability and automation across critical services. Responsibilities include monitoring production services, capacity planning, incident response and driving improvements to SLIs/SLOs. The role requires ...

Cybersecurity Analyst

Location
Birmingham, England, United Kingdom
engineers, mentoring junior team members. Partner with product, design and client stakeholders to translate business goals into technical solutions. Drive engineering excellence — testing, observability, performance and security best practice. Contribute to internal frameworks, design systems and reusable libraries across our consulting practice. What we're looking for 5+ years ...

VP, SRE: Architect Resilient, Scalable Systems

Location
Birmingham, England, United Kingdom
Vice President in Site Reliability Engineering (SRE) within Core Engineering at the Birmingham location. You will lead reliability efforts across distributed systems, driving SLOs, observability, and incident response while shaping scalable, automated platforms. This role emphasizes design reviews, reliability improvements, and reducing toil through automation, with a focus ...

SRE Associate: Build Reliable, Scalable Systems

Location
Birmingham, England, United Kingdom
Birmingham seeks an Associate Site Reliability Engineer to join Core Engineering. You will help build and operate high-performing, scalable systems, focusing on reliability, observability, and automation across critical services in a fast-paced financial environment. You will collaborate with engineering teams to reduce downtime, improve deployment speed, and implement ...

Senior IP Network Engineer & SRE Lead

Location
Birmingham, England, United Kingdom
engineering and SRE, leading complex fault resolution and end-to-end changes across BT’s fixed network infrastructure. You will champion reliability, automation and observability, delivering high-impact improvements while partnering with stakeholders. The role supports a 3 days in office, 2 days from home pattern across Ipswich, Birmingham ...

Agentic AI Engineer — AI Workflows & Orchestration

Location
Birmingham, England, United Kingdom
workflows within enterprise data environments in the UK. You will develop Python-based services, integrate with APIs and databases, and implement guardrails, testing and observability to production standards. Ideal candidates have 5–8 years in software/AI/ML roles, hands-on experience with GenAI applications beyond ...

SRE Engineer

Location
Birmingham, England, United Kingdom
month contract. Job Responsibilities/Objectives Implement SLOs, SLIs, error budgets and service health measures for critical services. Build and maintain observability patterns covering monitoring, logging, tracing, alert quality and dashboards. Support production readiness reviews, resilience reviews, capacity planning and operational acceptance. Automate repetitive operational tasks and reduce manual toil ...

Regional Lead Agentic Architect - Digital Engineer

Hiring Organisation
Hackajob Ltd
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent
trade-offs (autonomy vs control, cost vs performance, etc.) Nice-to-Have Skills Experience with tools like LangGraph, n8n, orchestration frameworks Knowledge of AI observability and evaluation frameworks Experience in sovereign or restricted deployment environments Background in ITSM or managed services automation Experience Profile 815+ years in architecture roles (solution ...

Infrastructure / DevOps Engineer

Location
Birmingham, England, United Kingdom
environments Define and implement infrastructure-as-code using Terraform, CDK, or equivalent Monitor platform health, define SLOs/SLAs, and build alerting and observability tooling Respond to and lead resolution of infrastructure-level incidents; drive post-mortems Harden infrastructure for HIPAA compliance — encryption, access controls, audit logging, and network security … Actions, Jenkins, or equivalent) Solid understanding of networking, security groups, load balancing, and DNS Experience with container orchestration (Docker, Kubernetes, or ECS) Familiarity with observability tooling (Datadog, Grafana, CloudWatch, or equivalent) Understanding of HIPAA infrastructure requirements (encryption at rest/in transit, audit trails, access controls) Nice to have Site ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
diverse engineering community. Design, build and maintain reliable, secure and scalable cloud-based infrastructure using infrastructure-as-code approaches. Enable teams to develop effective observability practices, including monitoring, logging, metrics and alerting that support proactive service management. Work with teams to define and embed Service Level Indicators (SLIs), Service Level … professionals, helping shape platform strategy, improve service reliability and support the delivery of critical digital services across government. The team is actively investing in observability, service-level management, platform automation, developer experience and cloud engineering. You'll join a culture that values collaboration, continuous learning and the freedom to explore ...

Forward Deployed AI Engineer

Location
Birmingham, England, United Kingdom
frontier models where reasoning genuinely earns it), prompt and context budgeting, caching. Cost is a non-negotiable metric on every engagement. Ship it properly. Observability and tracing, CI/CD, IaC, monitoring the client can actually operate, security and compliance review. Transfer knowledge deliberately. We don't run a long … tool exposure, auth patterns. Claude Code as an autonomous SDLC agent: sub‐agents, hooks, custom skills, agent marketplaces, context sharing across a team. LLM observability and tracing tooling (LangFuse, LangSmith, Arize, Braintrust or similar). Regulated‐environment delivery: HIPAA, GDPR, SOC 2, FCA/PRA, PII handling, data residency, guardrails ...