1,576 to 1,600 of 2,655 Remote/Hybrid Observability Jobs

Senior Cloud Architect (all genders)

Hiring Organisation
Lam Research
Location
Villach, Kärnten, Austria
Employment Type
Permanent
Salary
EUR Annual
guardrails. Developer Platform & Experience Build an internal developer platform (IDP) that abstracts complexity and provides self-service: project scaffolding, environment creation, CI/CD, observability, and secrets management. Curate platform catalogs/registries (Terraform module registry, container base images, reusable pipeline templates). Partner with product/engineering teams … Policy, OPA, Conftest) and continuous compliance. Build secure-by-default blueprints: network segmentation, private endpoints, encryption, vulnerability mgmt, and SBOM/SLSA practices. SRE, Observability, and Operations Embed SRE practices: SLOs, error budgets, incident response, postmortems, chaos/gamedays. Standardize observability: logs, metrics, traces, dashboards, and alerts (e.g., Azure Monitor ...

Observability Solutions Consultant - Client-Facing

Location
Greater London, England, United Kingdom
Itrs Insights is seeking a Professional Services Consultant for their London HQ. This role involves managing client projects and improving service offerings. Candidates should have at least 12 months of ITRS Geneos expertise, 3 years ...

Platform Engineer

Location
Greater London, England, United Kingdom
efficiently and securely. Working closely with software developers, architects, and delivery teams, you will help establish best practices around cloud infrastructure, CI/CD, observability, security, and application reliability. The role combines hands-on engineering with the opportunity to influence platform standards and development practices across multiple projects. Key Responsibilities … with modern front-end technologies, including React and Vite. Develop and improve CI/CD pipelines, automation, and developer tooling. Implement monitoring, logging, and observability solutions to improve system reliability and performance. Manage and optimise cloud infrastructure, containerised environments, and platform configurations. Ensure security and operational best practices are embedded ...

Senior DevOps Engineer - AVP

Location
Belfast City District, Northern Ireland, United Kingdom
operations on a global scale. In this role, you will apply deep technical expertise across CI/CD pipelines, container orchestration, cloud infrastructure, and observability to deliver resilient, high-quality software systems. Your work will directly shape how Citi's engineering teams build, ship, and monitor production services across … engineering teams. Build and manage containerized workloads on Kubernetes and OpenShift using Helm, ensuring systems are scalable, reliable, and production ready. Architect and maintain observability solutions — including log aggregation with Splunk and Elastic/Kibana, and metrics monitoring with Prometheus and Grafana — to give engineering teams real-time visibility into ...

Remote Staff Software Engineer - Databases SRE UK Remote

Location
Havant, Hampshire, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … self-directed or as the result of learnings from incidents, and may include improvements to monitoring, automation, increasing self-healing, auto-scaling, etc. Improve observability of customers within their environments Designing and implementing solutions to ensure reliability and scalability of our environments can meet rapidly increasing demands Develop fault-tolerant ...

Remote Staff Software Engineer - Databases SRE UK Remote

Location
Rhyl, Denbighshire, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … self-directed or as the result of learnings from incidents, and may include improvements to monitoring, automation, increasing self-healing, auto-scaling, etc. Improve observability of customers within their environments Designing and implementing solutions to ensure reliability and scalability of our environments can meet rapidly increasing demands Develop fault-tolerant ...

Remote Staff Software Engineer - Databases SRE UK Remote

Location
Alford, Aberdeenshire, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … self-directed or as the result of learnings from incidents, and may include improvements to monitoring, automation, increasing self-healing, auto-scaling, etc. Improve observability of customers within their environments Designing and implementing solutions to ensure reliability and scalability of our environments can meet rapidly increasing demands Develop fault-tolerant ...

Remote Staff Software Engineer - Databases SRE UK Remote

Location
St. Helens, Merseyside, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … self-directed or as the result of learnings from incidents, and may include improvements to monitoring, automation, increasing self-healing, auto-scaling, etc. Improve observability of customers within their environments Designing and implementing solutions to ensure reliability and scalability of our environments can meet rapidly increasing demands Develop fault-tolerant ...

Remote Staff Software Engineer - Databases SRE UK Remote

Location
Princes Risborough, Buckinghamshire, United Kingdom
Grafana Labs, the company behind the open observability cloud, is founded on the principles of open source, open standards, open ecosystems, and open culture. Grafana Cloud, our fully managed observability platform, is flexible and built for scale. With Grafana Cloud's actually useful AI, organizations can see, understand … self-directed or as the result of learnings from incidents, and may include improvements to monitoring, automation, increasing self-healing, auto-scaling, etc. Improve observability of customers within their environments Designing and implementing solutions to ensure reliability and scalability of our environments can meet rapidly increasing demands Develop fault-tolerant ...

Operations Engineering Lead

Hiring Organisation
Willis Towers Watson
Location
London, UK
Employment Type
Full-time
base platform operations team. You will help set direction for shared infrastructure and engineering operations - the production hosting, security perimeter, release pipelines, observability, and developer tooling that every engineering squad depends on to ship safely. This is a role that includes Line Management responsibility for your team … networking, scalability, performance, and cost-efficiency across production environmentsOversee Infrastructure as Code practices (e.g. Terraform) and ensure environments are consistent, auditable, and secureShape monitoring, observability and alerting across the engineering organisation (e.g. Datadog, CloudWatch) so issues are detected and resolved before customers are impacted4Lead incident management practices - on-call, triage ...

Operations Engineering Lead

Hiring Organisation
WTW
Location
Greater London, United Kingdom
Employment Type
Full Time
base platform operations team. You will help set direction for shared infrastructure and engineering operations - the production hosting, security perimeter, release pipelines, observability, and developer tooling that every engineering squad depends on to ship safely. This is a role that includes Line Management responsibility for your team … performance, and cost-efficiency across production environments Oversee Infrastructure as Code practices (e.g. Terraform) and ensure environments are consistent, auditable, and secure Shape monitoring, observability and alerting across the engineering organisation (e.g. Datadog, CloudWatch) so issues are detected and resolved before customers are impacted4 Lead incident management practices - on-call ...

Devops Engineer

Hiring Organisation
ISR RECRUITMENT LIMITED
Location
Nationwide, United Kingdom
Employment Type
Contract
Contract Rate
£475 - £500/day (Outside IR35)
SAML Application Technologies: React.js | Angular.js You will work extensively with Terraform, Kubernetes, GitHub Actions, Flux and GitOps, alongside cloud networking, security, identity, API management, observability and event-driven architectures. Role and Responsibilities: Deploy and manage infrastructure across multiple AWS cloud environments using Terraform and Infrastructure as Code (IaC). Deploy … manage and maintain cluster configurations. Securely configure and deploy AWS API Gateway and associated API capabilities. Design and implement effective operational monitoring, alerting and observability solutions. Design and support cloud networking across AWS, Azure and GCP, including VPCs, VNets, subnets, routing, load balancing and DNS. Deploy and operate event-driven ...

SENIOR BACKEND ENGINEER

Hiring Organisation
Widenet Consulting
Location
United States
Employment Type
Permanent
Salary
USD Hourly
vendor platforms Define and evolve API contracts (REST, GraphQL), JSON schemas, event contracts and canonical structures, with an emphasis on scalability, performance, and observability Contribute to architectural decisions, particularly around microservices, API-to-API, events, webservices/webhooks, in a real-world production context Develop a strong understanding … translate business needs into working software Contribute to the full service lifecycle-from development and CI/CD to infrastructure and deployment Champion observability practices including structured logging, monitoring, and instrumentation (Datadog or AWS Cloudwatch experience preferred) Drive engineering excellence by leading design reviews, contributing to post-mortems, and documenting ...

Principal Java Engineer

Location
Wallingford, England, United Kingdom
continuous improvement. Production systems are reliable, observable and operationally excellent Lead root cause analysis and resolution of complex production issues. Drive improvements in system observability, monitoring and operational performance. Ensure applications are designed and operated to meet reliability, availability and performance targets. Partner with Operations, DevOps and QA teams … Claude, Codex, Gitlab Duo, etc) REST APIs, OpenAPI, Microservices, Event-driven architecture (RabbitMQ) Containers, Docker, AWS, Linux CI/CD with GitLab Pipelines & Jenkins Observability: logging, metrics and monitoring MySQL, Apache Solr Front-end UI (e.g. Angular) Person Specification Strategic and systems-thinking mindset Excellent communication and stakeholder management skills ...

Senior Software Development Engineer

Location
Reading, England, United Kingdom
agile delivery environments (e.g. Scrum) Desirable Experience with cloud platforms (e.g. Azure) Experience with CI/CD pipelines and release processes Experience with observability tools (monitoring, logging, alerting) Experience in technical leadership or mentoring roles Experience working in safety- or regulation-driven environments Technology Stack .NET (latest versions) Kubernetes & Docker … Azure (SQL, CosmosDB, cloud services) PostgreSQL TypeScript/modern web frameworks (e.g. Vue.js) Observability tooling (e.g. Azure Monitor, Prometheus, Grafana) Azure DevOps/CI-CD pipelines Holidays: 25 days per annum + 8 days bank holidays (options to buy/sell days) 37.5 hour working week Pension – 4% employee ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
architecture and technical decision‐making as our platform evolves Champion engineering excellence through automated testing, CI/CD, code reviews and pair programming Use observability and monitoring to understand service health, troubleshoot complex distributed systems and improve reliability Mentor and support other engineers, sharing knowledge and helping raise technical capability … modern CI/CD practices Strong understanding of automated testing, code quality, security and software engineering best practices Experience troubleshooting distributed systems and using observability to improve performance and reliability Strong communication and collaboration skills, with experience contributing to technical decisions Experience mentoring and supporting other engineers through code reviews ...

Automation & Platform Engineer

Location
Newbury, England, United Kingdom
services, agent services, APIs, and microservices.* Implement infrastructure-as-code for platform environments and network automation resources.* Ensure automation platforms meet security, compliance, availability, observability, and operational resilience requirements.## **Your Profile*** Experience in network automation, platform engineering, DevOps, cloud engineering, or telecom automation.* Strong hands-on experience with Ansible, Terraform … vendor APIs.* API management and orchestration.* Intent-to-configuration workflows.* Data pipelines.* Vector databases and graph APIs.* MCP integration.* Security frameworks and compliance controls.* Observability and logging.**Preferred Certifications*** Kubernetes CKA/CKAD.* Terraform Associate.* Red Hat Ansible certification.* Google Cloud, AWS, or Azure certification.* Cisco, Juniper, Nokia, or Ericsson ...

Principal Engineer

Location
Greater London, England, United Kingdom
changes, unblocking delivery and coaching engineers through real work. Drive engineering excellence through test automation, trunk-based development, CI/CD optimisation, secure coding, observability, service-level objectives and effective production incident response. Embed AI-assisted engineering practices into day-to-day delivery, including drafting, review support, test generation … data and infrastructure. Deep experience with cloud-native platforms, Kubernetes or equivalent container orchestration, infrastructure-as-code, CI/CD pipelines, progressive delivery, observability and production operations. Confidence using AI-enabled engineering tools to improve quality, speed and traceability, while maintaining strong human oversight. Proven ability to build ...

Lead .Net Software Engineer

Location
Cheltenham, England, United Kingdom
technical risks, bottlenecks, dependencies, and opportunities for improvement Contribute to AWS‐based architecture and engineering practices, including environments using Lambda, ECS, and EC2 Improve observability, reliability, performance, security, and operational readiness across the platform Contribute to Jenkins pipelines and CI/CD practices to improve consistency and delivery efficiency Help … with DevOps practices and infrastructure‐aware development Experience with messaging, event streaming, or related event‐driven technologies Experience with Blazor and MudBlazor Familiarity with observability and monitoring tooling in distributed systems Experience defining governance, controls, or operating models for AI agents or AI‐enabled internal tools Experience working ...

Senior Backend Engineer

Location
Greater London, England, United Kingdom
Lead code reviews, mentor engineers, and ensure high engineering standards. Contribute to platform SDKs, internal libraries, and backend frameworks used across charters. Performance, Reliability & Observability Implement tracing, structured logging, metrics, dashboards, and alerting. Optimize services for latency, concurrency, throughput, and cost efficiency. Ensure system reliability through automated testing, load testing … Kubernetes, and cloud platforms (AWS or GCP). Strong CI/CD experience using GitHub Actions, Jenkins, Argo, or similar tools. Deep knowledge of observability practices — logs, metrics, traces, performance analysis. Proven ability to write clean, testable, well‐structured production code. Qualifications – Nice to Have Experience with streaming or queuing ...

Data Platform Engineer

Hiring Organisation
MONY Group
Location
London, UK
Employment Type
Full-time
personalisation. Stay close to the business context and apply software engineering practices to solve specific data problems. Improve monitoring, alerting, data quality checks and observability so issues are detected and understood quickly. Improve engineering workflows with AI and automationIdentify repeated or high-friction engineering tasks and turn them into reliable … improving cloud-based systems. Familiarity with infrastructure-as-code, CI/CD, version control and automated testing. Ability to reason about reliability, security, observability and operational support. Experience working with technical and non-technical stakeholders and communicating clearly. Curiosity about AI-assisted engineering and automation, with an interest in applying ...

The Core Engineering - Site Reliability Engineering - Associate - Birmingham

Location
Birmingham, England, United Kingdom
operational resilience. Practice sustainable incident management through clear escalation, effective remediation, and a blameless postmortem culture. Identify and implement improvements to system behavior, controls, observability, and monitoring tools. Define and maintain service level indicators (SLIs), service level objectives (SLOs), and error budgets to quantify and manage service reliability. Engineer automation … development lifecycle concepts, and developing applications in a Linux environment. Strong understanding of algorithms, data structures, software design, and distributed systems fundamentals. Experience with observability platforms, including distributed tracing, logging, metrics, and tools such as Prometheus, Grafana, ELK, or OpenTelemetry. Experience with site reliability engineering practices, relational databases, Hadoop ...

Lead Software Engineer - AI-Native Applications

Hiring Organisation
IFS
Location
London, UK
Employment Type
Full-time
technologies such as Kubernetes, Kafka, Redpanda, PostgreSQL and MongoDB.Experience working with AWS and/or Azure. Strong understanding of CI/CD, automated testing, observability and production operations. Experience designing secure software, authentication and authorisation mechanisms and applying DevSecOps principles. Strong analytical and problem-solving skills with the ability … Experience building Enterprise SaaS or ERP products. Experience working with Model Context Protocol (MCP) or similar AI integration standards. Experience with vector databases, AI observability or AI evaluation frameworks. Experience designing and delivering enterprise-scale AI platforms or AI-powered products. QualificationsA degree in Computer Science, Software Engineering or Information ...

Data Engineer, Vice President

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent, Work From Home
datasets, and extensible pipelines that support multiple Company Intelligence products and advanced analytics use cases. Ensure high standards of data quality, governance, lineage, and observability, proactively managing operational and compliance risks across enterprise-grade data products. Partner effectively with product, analytics, and business leaders, developing a deep understanding of strategic … control, testing, CI/CD). Experience building and operating data pipelines using workflow orchestration frameworks (e.g. Apache Airflow), with a focus on reliability, observability, dependency management, and operational resilience. Experience designing and operating cloud-native data platforms (AWS or Azure preferred) and enterprise data warehouses (Snowflake preferred), including performance ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
system integrations Design reliable APIs, data models, caching, synchronization, and recovery paths using PostgreSQL, Valkey, REST, and service integrations across the Cloudbeds platform Own observability, automated testing, release quality, performance, security, and production support for customer‐critical workflows Own changes across the POS application, backend services, infrastructure, and deployment configuration … asynchronous or event‐driven workflows Experience owning cloud‐native delivery and operations, including Terraform, Docker, Kubernetes, Helm, Argo CD, AWS, CI/CD, observability, and production incident response Experience building point‐of‐sale or adjacent real‐time transactional systems where money, inventory, and people meet at the moment of service ...