1,851 to 1,875 of 5,503 Permanent Observability Jobs

Vice President, Production Services Application Support

Location
Manchester, England, United Kingdom
Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe … Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe ...

Full Stack Developer - React & Node.js

Hiring Organisation
Shree Narayani Networking Solutions Pvt Ltd
Location
San Mateo, California, United States
Employment Type
Permanent
Salary
USD Annual
party systems. Comfort with event-driven or streaming systems feeding real-time consumers, and the reliability mindset that comes with owning a live service (observability, retries, graceful degradation). Ability to own a service across the stack - stand it up, deploy it, monitor it - without needing hand-holding on infra. … modeling - designing canonical schemas and reconciling inconsistent inputs across sources. Real-time/event-driven systems - message queues or streaming (e.g., Kafka), and the observability tooling to run them. Cloud & operations - deploying and monitoring services in a cloud environment (AWS/GCP), containers, CI/CD. Strong signals (nice ...

Staff Analytics Platform Engineer

Location
Greater London, England, United Kingdom
that improve performance, developer experience, cost efficiency, or operational maturity. Owning and evolving core platform components, including CI/CD, testing strategies, environment management, observability, and infrastructure as code. Acting as the technical escalation point for complex, cross‐cutting platform issues and guiding teams toward robust, scalable solutions. Driving Snowflake … performance and cost optimisation, informed by real workloads and modelling patterns. Implementing and maturing data SLAs/SLOs, data observability, lineage, and quality frameworks to ensure trusted analytics at scale. Collaborating with data product and engineering teams to enable safe, scalable ingestion and well‐defined data contracts. Influencing how teams ...

Lead SRE - AWS Platform

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
your team to identify comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets Design and implement observability frameworks and alerting strategies, including white and black box monitoring, service level objective-based alerting, and telemetry collection to ensure proactive detection and response Serve … resiliency best practices Fluency in at least one programming language such as Python, Java/Spring Boot, or .NET Proficient knowledge and experience in observability, including white and black box monitoring, service level objective alerting, and telemetry collection across large-scale production environments Proficiency with continuous integration and continuous delivery ...

Site Reliability Engineer

Location
Fenny Stratford, England, United Kingdom
hands‐on role in ensuring it is reliable, scalable, and observable. You will help establish and mature SRE practices, focusing on: Monitoring and observability Reliability testing and capacity planning Toil reduction We offer a hybrid working arrangement with one day per week in our Milton Keynes office. Key Responsibilities: Support … Build dashboards, alerts, and runbooks to improve visibility Automate repetitive tasks to reduce operational toil Collaborate with cross-functional teams to enhance reliability and observability Support performance testing and capacity planning Proactively identify and prioritise reliability improvements Experience & Skills Required: Hands‐on experience with Azure Monitoring (Application Insights, Alerts, Action ...

Full Stack Developer - Advanced

Hiring Organisation
BC Forward
Location
Columbus, Ohio, United States
Employment Type
Permanent
Salary
USD 80 Hourly
Advanced to join our team. The ideal candidate will have strong experience in Java/Spring Boot, React/TypeScript, AWS, Kafka, SQL, and observability and a proven ability to design, build, secure, and operate scalable full-stack systems in production. Responsibilities: Design, implement, and operate backend services using Java …/Aurora, S3, SQS/SNS, and CloudWatch. Create infrastructure as code using Terraform or CloudFormation/CDK with environment promotion. Establish observability with structured logging, metrics, tracing, dashboards, and actionable alarms. Drive quality through automated testing, performance testing, and secure SDLC practices. Contribute to runbooks, on-call readiness ...

Lead Cloud Site Reliability Engineer

Location
Halifax, England, United Kingdom
deliver secure, resilient and scalable services for millions of customers. We're looking for a Site Reliability Engineer Lead to help strengthen reliability, observability and operational excellence across our Azure and Google Cloud Platform (GCP) environments. You'll lead a team of Site Reliability Engineers, helping to establish engineering standards … supports learning, collaboration and continuous improvement. Partner with Product Owners, Engineering Leads and platform teams to balance reliability, operational resilience and feature delivery. Use observability data, platform metrics and service insights to identify improvement opportunities and reduce operational risk. Lead incident and problem management activities, promoting effective root cause analysis ...

Site Reliability Engineer

Hiring Organisation
Connells Limited
Location
Milton Keynes, Buckinghamshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
hands-on role in ensuring it is reliable, scalable, and observable. You will help establish and mature SRE practices, focusing on: Monitoring and observability Incident response Post-incident review Reliability testing and capacity planning Toil reduction Enabling development velocity We offer a hybrid working arrangement with one day per week … Build dashboards, alerts, and runbooks to improve visibility Automate repetitive tasks to reduce operational toil Collaborate with cross-functional teams to enhance reliability and observability Support performance testing and capacity planning Proactively identify and prioritise reliability improvements Experience & Skills Required: Hands-on experience with Azure Monitoring (Application Insights, Alerts, Action ...

AI Engineer

Location
Greater London, England, United Kingdom
products. As a Senior/Lead AI Engineer, you will own and drive the development of our core agentic frameworks, evaluation pipelines, and observability tooling to ensure they operate with safety, trust, and intelligence at scale. You won't just be integrating AI into a product, you will lead … CrewAI is a plus. Production Python engineering : Clean, modular, testable, maintainable code — you care about system reliability as much as model output. Evaluation & observability : Practical experience instrumenting tracing (LangSmith, Arize) and building CI/CD pipelines built specifically for LLMs. Engineering foundations : Distributed systems, scalable data pipelines (Kafka ...

Security Platform Engineer

Location
Farnborough, England, United Kingdom
agents Hands‐on experience with Kubernetes Experience managing and administering SIEM tooling (e.g. Splunk) Experience deploying vulnerability scanning and analysis tooling (e.g. Nessus) Kubernetes observability (e.g. fluentbit) Familiarity with container security principles and tools Scripting or automation skills (e.g. Python, Bash, or similar) Understanding of security frameworks and best practices … Knowledge of configuring SIEM tooling Basic understanding of threat frameworks, such as ATT&CK Experience with additional SIEM or observability platforms Experience with Microsoft Defender Experience with DevSecOps practices and pipeline security Certifications such as CKA, CKAD, CISSP, CEH, or similar Exposure to threat modelling and security architecture design Just ...

Network Automation Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
Python, Ansible, Terraform and Jinja2Integrating network automation into CI/CD pipelines for reliable, repeatable deploymentsCreating APIs and self-service tooling for engineering teamsImplementing observability and telemetry solutions for performance and reliabilityPartnering with network, platform and security teams to deliver resilient, scalable systemsContributing to incident response and production reliabilityOn-call … tools such as Ansible, Terraform and Jinja2; also must have experience leveraging AI tools, such as Claude CodeFamiliarity with Docker and KubernetesExposure to monitoring, observability or telemetry in distributed systemsPragmatic problem solver who can operate in ambiguity and take ownershipComfortable working in collaborative, fast-paced engineering teamsDeep understanding of networking ...

Senior Software Engineer

Hiring Organisation
Ask4.com
Location
Sheffield, South Yorkshire, Yorkshire, United Kingdom
Employment Type
Permanent
Salary
£60,000
Kubernetes configurations for production and non-production environments Integrate with network management systems, message brokers, and third-party APIs Instrument and monitor applications using observability tooling (metrics, logs, and traces Grafana, Prometheus, or similar) Provide technical mentorship to mid-level engineers and act as an escalation point for complex problems … Experience with message brokers and event-driven systems (NATS, RabbitMQ, Kafka, or similar) Exposure to OpenWiFi, OpenWrt or similar open-source network controller frameworks Observability experience. Use of metrics, logs, and traces using tools such as Grafana, Prometheus and Sentry Experience with AI agent development or LLM integration Agile/ ...

Database Reliability Engineer

Location
Manchester, England, United Kingdom
Cloud Portability: Use CNPG and cloud-native patterns to ensure our database layer remains provider-agnostic, allowing seamless deployment across AWS and GCP Evolve Observability & Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will ensure we have the visibility to detect performance regressions and health … Cloud Portability: Use CNPG and cloud-native patterns to ensure our database layer remains provider-agnostic, allowing seamless deployment across AWS and GCP Evolve Observability & Monitoring: Build deep, proactive monitoring and alerting for our global database fleet. You will ensure we have the visibility to detect performance regressions and health ...

Network Automation Engineer

Location
Greater London, England, United Kingdom
Jinja2 Integrating network automation into CI/CD pipelines for reliable, repeatable deployments Creating APIs and self‐service tooling for engineering teams Implementing observability and telemetry solutions for performance and reliability Partnering with network, platform and security teams to deliver resilient, scalable systems Contributing to incident response and production reliability … Ansible, Terraform and Jinja2; also must have experience leveraging AI tools, such as Claude Code Familiarity with Docker and Kubernetes Exposure to monitoring, observability or telemetry in distributed systems Pragmatic problem solver who can operate in ambiguity and take ownership Comfortable working in collaborative, fast‐paced engineering teams Deep understanding ...

Senior Software Engineer-AI

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent
operation with limited supervision Hands-on experience with cloud-native technologies, serverless applications, event-driven architectures, data pipelines, relational and NoSQL databases, vector databases, observability tooling, and automated deployment pipelines Solid understanding of algorithms, data structures, scalability, reliability, performance optimization, security best practices, and engineering trade-offs Experience mentoring engineers … technical designs, participate in design reviews, and identify risks, constraints, trade-offs, and alternative approaches Maintain engineering excellence through automated testing, code reviews, observability, monitoring, alerting, operational readiness, and participation in on-call support Apply machine learning operations practices, including prompt versioning, automated evaluation, deployment pipelines, monitoring, and production issue ...

Developer – Scala

Location
Newcastle upon Tyne, England, United Kingdom
frontend engineers, QA, product owners, solution designers, and other backend developers to deliver high-quality product increments. Support production stability by investigating issues, improving observability, and continuously reducing technical debt. Required Skills and Experience: Professional backend development experience, ideally in enterprise, SaaS, or cloud-based product environments. Strong hands … problems. Nice to Have: Experience with Squeryl, Doobie, or similar Scala data access libraries. Familiarity with Grafana or Kibana dashboards, alerting, logging, and production observability practices. Experience with large distributed systems, horizontal scaling, resilient service design, or high-throughput Play/Pekko applications. Previous experience in Payroll ...

Tivoli Netcool OMNIbus SME

Hiring Organisation
Proactive Appointments
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£750.00 - £815.00 per day
effective monitoring coverage. Define platform standards, monitoring policies and best practices. Maintain technical documentation, architecture diagrams and support procedures. Contribute to the monitoring and observability roadmap and identify opportunities for platform modernisation. Provide technical mentoring and knowledge sharing across engineering and operational teams. Essential Experience Significant hands-on experience administering … integration technologies. Understanding of monitoring across AWS, Azure and Google Cloud . Experience with container and Kubernetes monitoring . Knowledge of enterprise observability frameworks and modern SRE practices . Tivoli Netcool OMNIbus SME Due to the volume of applications received for positions, it will not be possible to respond ...

Vice President, Production Services Application Support

Hiring Organisation
Hackajob Ltd
Location
Manchester, North West, United Kingdom
Employment Type
Permanent
Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe … resolve complex technical issues under pressure. Deep understanding of enterprise application architecture, distributed systems, cloud technologies, middleware, databases, and infrastructure components. Experience with monitoring, observability, automation, and operational tooling used to support highly available production platforms. Strong analytical and problem-solving skills with the ability to identify root causes ...

Principal Agentic Architect (all genders)

Hiring Organisation
Lam Research
Location
Villach, Kärnten, Austria
Employment Type
Permanent
Salary
EUR Annual
Architect, you will define the architecture that enables AI agents to reason, automate, and operate at enterprise scale while ensuring security, governance, reliability, and observability remain foundational. You will establish the long-term vision for how AI agents, human operators, and digital platforms work together to create a highly automated … patterns, and governance frameworks for AI agents and multi-agent systems. Design the integration architecture connecting agents to enterprise systems including ServiceNow, cloud platforms, observability platforms, CMDB, developer platforms, and security tooling. Develop reference architectures for AI-powered incident response, service management, platform operations, FinOps, cybersecurity, and disaster recovery. Partner ...

Senior Software Engineer

Hiring Organisation
The Portfolio Group
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£90000/annum
Establish automated testing across backend and frontend applications, including unit, contract and end-to-end testing. Work with the platform engineering team on deployment, observability, logging, tracing and operational readiness. Act as the technical owner for the application and integration layer, making and documenting key architectural decisions. Provide technical guidance … such as Lambda, ECS, API Gateway, S3, CloudFront, Cognito and IAM. Experience designing and operating distributed or event-driven systems. A strong understanding of observability, testing and CI/CD practices. Experience working with data platforms or stores such as MongoDB, OpenSearch or Databricks. Experience integrating internal systems and third ...

Forward Deployed Agentic AI Engineer

Hiring Organisation
WTW
Location
Greater London, United Kingdom
Employment Type
Full Time
solutions that solve real-world business challenges. You will bring deep expertise across modern full-stack technologies, distributed systems, cloud-native architectures, and observability, combined with hands-on experience developing enterprise-grade AI applications. You will design, build, and deploy intelligent AI agents, copilots, and automation solutions using Anthropic Claude … orchestration, evaluation loops, and human-in-the-loop controls. Enterprise integration: Integrate AI solutions with enterprise systems, APIs, data platforms, document repositories, workflow tools, observability platforms, and identity and access management services. Production engineering: Ensure AI solutions meet enterprise standards for reliability, scalability, latency, maintainability, cost control, logging, monitoring ...

Senior Software Engineer

Hiring Organisation
The Portfolio Group
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£90,000 per annum
Establish automated testing across backend and frontend applications, including unit, contract and end-to-end testing. Work with the platform engineering team on deployment, observability, logging, tracing and operational readiness. Act as the technical owner for the application and integration layer, making and documenting key architectural decisions. Provide technical guidance … such as Lambda, ECS, API Gateway, S3, CloudFront, Cognito and IAM. Experience designing and operating distributed or event-driven systems. A strong understanding of observability, testing and CI/CD practices. Experience working with data platforms or stores such as MongoDB, OpenSearch or Databricks. Experience integrating internal systems and third ...

Senior Backend Engineer - Asset Sales

Location
Greater London, England, United Kingdom
Modern C# stack : Distributed C# and .NET microservices Cloud & orchestration : Hosted on Azure using Kubernetes Architecture : Event-driven, supporting products used at significant scale Observability : Grafana, Azure Application Insights, logs, traces, and metrics AI tooling : Claude and other AI tools used throughout the engineering workflow — design exploration, code generation … want engineers who tinker — experimenting with new tools, agents, and workflows, and sharing what works Guardrails as we accelerate : Automated tests, SLOs, alerting, observability, and deployment safeguards around everything we ship Own it beyond the pull request : Design for idempotency, retries, out-of-order events, and failure modes, and know ...

Embedded Software Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, UK
Employment Type
Full-time
security best practices across the device lifecycleEdge and cloud integration: integrate devices with cloud IoT platforms and backend services; improve telemetry, health monitoring and observability; support reliable field operation and debug fleet issuesCompliance and testing: support testing and validation for safety, EMC, radio and related requirements; prepare test plans …/CD workflows for embedded softwarePractical debugging experience using lab and software tools such as logic analysers, protocol analysers, network sniffers or observability platformsStrong problem-solving skills and the ability to work across hardware and software boundariesBonus: low-power wireless (Thread, Zigbee, BLE); cloud IoT services on AWS, Azure ...

AI Platform Support Engineer (EMEA)

Location
Greater London, England, United Kingdom
combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in. We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle … post incident reviews and operational improvements Build internal tooling, automation, documentation, and runbooks Partner closely with infrastructure, networking, and platform engineering teams Help improve observability, operational visibility, and troubleshooting workflows Improve the customer experience through better processes and technical guidance What This Role Is Not This is not a traditional ...