1,726 to 1,750 of 4,833 Observability Jobs

Tivoli Netcool OMNIbus SME - 11850CF

Hiring Organisation
Proactive Appointments
Location
London, UK
Employment Type
Full-time
effective monitoring coverage. Define platform standards, monitoring policies and best practices. Maintain technical documentation, architecture diagrams and support procedures. Contribute to the monitoring and observability roadmap and identify opportunities for platform modernisation. Provide technical mentoring and knowledge sharing across engineering and operational teams. Significant hands-on experience administering and supporting … APIs, JSON, XML and integration technologies. Understanding of monitoring across AWS, Azure and Google Cloud. Experience with container and Kubernetes monitoring. Knowledge of enterprise observability frameworks and modern SRE practices. Tivoli Netcool OMNIbus SMEDue to the volume of applications received for positions, it will not be possible to respond ...

Senior Software Engineer

Hiring Organisation
The Portfolio Group
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£90,000 per annum
Establish automated testing across backend and frontend applications, including unit, contract and end-to-end testing. Work with the platform engineering team on deployment, observability, logging, tracing and operational readiness. Act as the technical owner for the application and integration layer, making and documenting key architectural decisions. Provide technical guidance … such as Lambda, ECS, API Gateway, S3, CloudFront, Cognito and IAM. Experience designing and operating distributed or event-driven systems. A strong understanding of observability, testing and CI/CD practices. Experience working with data platforms or stores such as MongoDB, OpenSearch or Databricks. Experience integrating internal systems and third ...

Senior Backend Engineer - Asset Sales

Location
Greater London, England, United Kingdom
Modern C# stack : Distributed C# and .NET microservices Cloud & orchestration : Hosted on Azure using Kubernetes Architecture : Event-driven, supporting products used at significant scale Observability : Grafana, Azure Application Insights, logs, traces, and metrics AI tooling : Claude and other AI tools used throughout the engineering workflow — design exploration, code generation … want engineers who tinker — experimenting with new tools, agents, and workflows, and sharing what works Guardrails as we accelerate : Automated tests, SLOs, alerting, observability, and deployment safeguards around everything we ship Own it beyond the pull request : Design for idempotency, retries, out-of-order events, and failure modes, and know ...

Embedded Software Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, UK
Employment Type
Full-time
security best practices across the device lifecycleEdge and cloud integration: integrate devices with cloud IoT platforms and backend services; improve telemetry, health monitoring and observability; support reliable field operation and debug fleet issuesCompliance and testing: support testing and validation for safety, EMC, radio and related requirements; prepare test plans …/CD workflows for embedded softwarePractical debugging experience using lab and software tools such as logic analysers, protocol analysers, network sniffers or observability platformsStrong problem-solving skills and the ability to work across hardware and software boundariesBonus: low-power wireless (Thread, Zigbee, BLE); cloud IoT services on AWS, Azure ...

AI Platform Support Engineer (EMEA)

Location
Greater London, England, United Kingdom
combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in. We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle … post incident reviews and operational improvements Build internal tooling, automation, documentation, and runbooks Partner closely with infrastructure, networking, and platform engineering teams Help improve observability, operational visibility, and troubleshooting workflows Improve the customer experience through better processes and technical guidance What This Role Is Not This is not a traditional ...

Senior Associate, Full-Stack Engineer Opportunities

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
ways: Design, build, and maintain backend services, batches and APIs, contributing to UI components as needed. Own end-to-end delivery: implementation, testing, deployment, observability, and reliability. Write clean, well-tested code; participate in code reviews and continuous improvement. Collaborate with product, design, and operations to translate business needs into … microservices Proficiency in Java with Spring. Experience with CI/CD, automated testing (JUnit/Spock), and containers (Docker). Familiarity with microservices, observability/telemetry (e.g., Splunk, AppDynamics), and cloud deployments. Curiosity to understand the business domain and translate product strategy into technical solutions. Working knowledge of Groovy with ...

Embedded Software Engineer

Location
Greater London, England, United Kingdom
best practices across the device lifecycle Edge and cloud integration: integrate devices with cloud IoT platforms and backend services; improve telemetry, health monitoring and observability; support reliable field operation and debug fleet issues Compliance and testing: support testing and validation for safety, EMC, radio and related requirements; prepare test plans …/CD workflows for embedded software Practical debugging experience using lab and software tools such as logic analysers, protocol analysers, network sniffers or observability platforms Strong problem-solving skills and the ability to work across hardware and software boundaries Bonus: low-power wireless (Thread, Zigbee, BLE); cloud IoT services ...

Lead Software Engineer - Cloud

Location
Glasgow, Scotland, United Kingdom
usability, and self-service for engineering consumers Create and maintain delivery workflows that standardize engineering practices and reduce operational toil across the platform Improve observability and operational readiness through monitoring, logging, tracing, alerting, runbook development, and on-call practices Partner with security and risk stakeholders to implement secure-by-default … Familiarity with infrastructure-as-code and automation practices, including tools such as Terraform, Helm, Kustomize, Argo CD, Flux, or continuous integration systems Experience with observability stacks and site reliability engineering practices, including service level indicators and objectives, incident response, and post-incident reviews Exposure to regulated environments and implementing security ...

Staff AI Engineer - EU

Hiring Organisation
Typeform
Location
United Kingdom
Salary
£ 70 K
questions, and turn responses into useful insights.The team owns the journey from experimentation through to production. This includes AI application development, evaluation, infrastructure, deployment, observability, reliability, and performance.You will work closely with Product Managers, Software Engineers, Data Scientists, Data Engineers, and Analytics teams to turn promising AI ideas into secure … difficult technical decisions and make trade-offs explicit.Mentor engineers and support other technical leads in growing their ownership and judgement.Improve engineering practices across testing, observability, security, incident response, and deployment.Build alignment around technical decisions through clear proposals, constructive discussion, and evidence.Evaluate relevant AI research and emerging tools, and help teams ...

Senior Full Stack Engineer - Lyst Shop (11 Month FTC - Maternity Cover)

Location
Greater London, England, United Kingdom
rely heavily on experimentation to validate ideas and guide decisions. Technical Excellence: You will help maintain a high bar for code quality, testing, observability, and system reliability. You’ll contribute to architectural discussions, improve developer experience, and proactively address technical debt where needed. Team Contribution: You will mentor and support … working relationships across Product, Design, QA, Analytics, and Engineering teams while actively participating in team ceremonies and technical discussions. Technical Impact: Improve the stability, observability, and maintainability of our systems through better monitoring, resilient code, and thoughtful testing practices. Growth & Ownership: Gain confidence working across our platform and infrastructure while ...

Staff AI Engineer - EU

Hiring Organisation
Typeform
Location
United Kingdom, UK
Employment Type
Full-time
turn responses into useful insights. The team owns the journey from experimentation through to production. This includes AI application development, evaluation, infrastructure, deployment, observability, reliability, and performance. You will work closely with Product Managers, Software Engineers, Data Scientists, Data Engineers, and Analytics teams to turn promising AI ideas into secure … decisions and make trade-offs explicit. Mentor engineers and support other technical leads in growing their ownership and judgement. Improve engineering practices across testing, observability, security, incident response, and deployment. Build alignment around technical decisions through clear proposals, constructive discussion, and evidence. Evaluate relevant AI research and emerging tools ...

Infrastructure Engineer

Location
Greater London, England, United Kingdom
tenant environments for enterprise customers. Developer Velocity: Architect CI/CD pipelines (GitHub Actions, Docker) that allow our team to ship safely and quickly. Observability & Reliability: Instrument the stack with OpenTelemetry and Datadog to ensure we detect issues before our users do. AI Performance Tuning: Tune autoscaling and network routing … Haves: Experience with ECS, container orchestration, and distributed task queues (Celery/SQS). Strong Python skills. Familiarity with the Datadog/Grafana observability stack. A background in offensive security, CTFs, or AI/ML infrastructure. What We Offer Competitive Salary + Significant Equity: We want you to have true ...

Security Platform Engineer

Hiring Organisation
IBM SIXworks Limited
Location
Farnborough, Hampshire, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
agents Hands-on experience with Kubernetes Experience managing and administering SIEM tooling (e.g. Splunk) Experience deploying vulnerability scanning and analysis tooling (e.g. Nessus) Kubernetes observability (e.g. fluentbit) Familiarity with container security principles and tools Scripting or automation skills (e.g. Python, Bash, or similar) Understanding of security frameworks and best practices … Knowledge of configuring SIEM tooling. Basic understanding of threat frameworks, such as ATT&CK. Experience with additional SIEM or observability platforms Experience with Microsoft Defender Experience with DevSecOps practices and pipeline security Certifications such as CKA, CKAD, CISSP, CEH, or similar Exposure to threat modelling and security architecture design Just ...

Software Engineer, Agentic Systems Engineering - ThousandEyes

Hiring Organisation
CISCO Systems
Location
London, UK
Employment Type
Full-time
experiences. ThousandEyes is deeply integrated across the Cisco technology portfolio and beyond, delivering AI-powered assurance insights within Cisco's Networking, Security, Collaboration, and Observability products. Our team develops internal AI-powered tools that help software engineers work more effectively across the development lifecycle. We build agentic workflows, reusable skills … integrations, or engineering productivity tooling. Knowledge of Kotlin, Java, Go or TypeScript. Understanding of Kubernetes, containers, and modern cloud deployment practices. Experience with observability, security, privacy, and responsible AI practices. Familiarity with relational, document-based, or vector datastores. Why Cisco? At Cisco, we're revolutionizing how data and infrastructure connect ...

Senior AI / Agentic Engineer London

Location
Greater London, England, United Kingdom
Agentic ecosystem, responsible for the high-level design choices that define how agents run at PhysicsX. You will cover topics such as: Agent Observability: Own the implementation to enforce deep tracing, granular cost tracking, and observability across the lifecycle. Agent Deployment: Deliver an intuitive deployment lifecycle which simplifies questions around … behalf of users in a regulated enterprise environment. The Tech Stack Core Platform: Python (Primary), Go or TypeScript (Secondary), Kubernetes, Docker, Terraform. Observability & Evals: OTel, LangSmith, Arize, Braintrust. Who You Are An Architect at Heart: You have strong, reasoned opinions on Durable Execution vs. Standard Async, Vector Search vs. Keyword ...

Security Platform Engineer: Build Secure Infra & CI/CD

Location
Farnborough, England, United Kingdom
agents Hands-on experience with Kubernetes Experience managing and administering SIEM tooling (e.g. Splunk) Experience deploying vulnerability scanning and analysis tooling (e.g. Nessus) Kubernetes observability (e.g. fluentbit) Familiarity with container security principles and tools Scripting or automation skills (e.g. Python, Bash, or similar) Understanding of security frameworks and best practices … Knowledge of configuring SIEM tooling. Basic understanding of threat frameworks, such as ATT&CK. Experience with additional SIEM or observability platforms Experience with Microsoft Defender Experience with DevSecOps practices and pipeline security Certifications such as CKA, CKAD, CISSP, CEH, or similar Exposure to threat modelling and security architecture design Just ...

Senior AI Engineer - Agentic AI

Location
Greater London, England, United Kingdom
that allow complex AI workflows to operate securely and efficiently at scale. You will be responsible for developing advanced orchestration capabilities, implementing evaluation and observability tooling and embedding enterprise controls for compliance and safety. If you are passionate about innovating with AI in real-world applications and scaling intelligent systems … ensuring graceful degradation and retries Apply enterprise security and governance practices including RBAC, prompt safety checks, traceability and secrets management Implement evaluation pipelines and observability frameworks using tools such as Langfuse, Arize or OpenTelemetry Contribute to architectural design decisions, code reviews and engineering standards for platform development Requirements Bachelor ...

Staff Python Engineer (ML)

Location
City Of London, England, United Kingdom
apps in a service architecture. Furthering Developer Experience (DevEx) by mentoring others in writing code that is intuitive, clear, and easy to test Developing observability for new and existing ML applications and GenAI/LLM integrations , making use of the Grafana Stack (Prometheus, Loki, Tempo) Develop integrations and services that … Backend-Engineering Experience owning projects from start to finish, including speccing, architecture, development, testing, deployment, release and monitoring Strong skills in building maintainable tests, observability and tracing systems. Knowledge of best practices for performance optimisation, memory management. Familiarity with Kubernetes , Docker and other cloud infrastructure, ops and containerised tools. Strong ...

Junior DevOps Engineer

Location
Greater London, England, United Kingdom
/CD processes, automate manual processes to create a self-service environment for our developers, and maintain platform uptime SLAs by improving our observability stack. We use infrastructure-as-code to maintain our platform on AWS, so familiarity with common AWS Services (RDS, S3, ECS, EC2 etc.) and Terraform …/CD pipelines through Jenkins/Github Actions and other technology Support the Development and AI Engineering Teams - help troubleshoot their issues Bring observability through dashboards, alerting and log aggregation Run incident analysis and post mortems Environment management: ephe...dev/staging/prod What you’ll bring A strong foundation ...

Software Engineer / AI Engineer

Location
Greater London, England, United Kingdom
within a defined problem, building and testing tool use, retrieval pipelines and agent workflows, integrating AI capabilities into enterprise systems, and contributing to evaluation, observability and guardrails. You will hold a high bar on code quality, flag risks and blockers early, and work alongside host‐function stakeholders to make sure … agentic AI solutions to production standard within a defined technical approach. Implement and test tool use, retrieval pipelines, and agent workflows. Contribute to evaluation, observability and guardrails for agentic systems. Integrate AI capabilities into existing enterprise workflows and systems. Maintain high code quality and documentation so patterns can be reused. ...

Mid-Level Data Engineer (Python/ AWS)

Location
Belfast, Northern Ireland, United Kingdom
consumers Contribute to cloud-based data platform development in AWS Support lightweight frontend work (React/TypeScript) for data-focused tools where needed Implement observability practices (logging, monitoring, tracing) across data pipelines and APIs Improve reliability, performance, and failure handling across the platform Collaborate with engineers and analysts to deliver … skills and experience working in cross-functional teams Nice to Have Experience with TypeScript or JavaScript Experience contributing to frontend applications (React) Familiarity with observability tooling (logging, monitoring, tracing) Experience with data modelling and metadata management Exposure to CI/CD practices for data pipelines and services Experience working ...

Full Stack Lead Software Engineer

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
design and development of React-based micro-frontend applications and shared UI frameworks. Own and drive system design and architecture decisions (scalability, resiliency, observability, security, and maintainability).Design and build container-based, scalable backend services using Python (Fast API) and modern API patterns. Develops secure and high-quality production code … capabilities, and skillsDesign and development of React-based micro-frontend applications and shared UI frameworks. Experience in system design and architecture decisions (scalability, resiliency, observability, security, and maintainability).Design and build container-based, scalable backend services using Python (FastAPI) and modern API patterns. Build and integrate services with ...

Member of Technical Staff (Platform Leaning)

Location
Greater London, England, United Kingdom
light up rather than back away: complex networking scenarios such as site-to-site VPN and private connectivity, multi-cloud and multi-region deployments, observability, and the machinery that makes Tessl straightforward to deploy, operate and support wherever it needs to run. We expect you to be relentlessly AI native … first year, you're shipping product features end to end like any Tessl engineer. When infrastructure work arises, from deployment architecture to networking to observability, you're the one who takes it on and lands it well, with support where you need it. Our deployment and operational practices are stronger ...

Jobshare - Sr Lead Software Engineer - Site Reliability Engineer, Python & Infrastructure management - Part time/Jobshare

Location
Greater London, England, United Kingdom
impact). Champions AI adoption and deliver AI-enabled capabilities to reduce operational toil and improve speed/quality of response. Sets direction for observability across logs/metrics/traces, including instrumentation standards, golden signals, and end-user journey monitoring. Improves alert quality and routing: reduce false positives, improve … design, build, and deliver automation and reliability solutions. Fluency & expertise in Python Deep practical knowledge of: SLOs/SLIs, error budgets, incident management, postmortems, observability design across metrics/logs/traces and distributed systems troubleshooting, resilience engineering, performance/capacity management, and change risk reduction. Proficiency and experience with ...

Principal Platform Engineer (12 Month FTC)

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
delivery, operational processes, and platform management. Establish and promote platform standards, engineering patterns, and best practices across engineering teams. Improve the reliability, scalability, performance, observability, and operational efficiency of the platform. Identify opportunities to reduce engineering friction, eliminate repetitive work, and accelerate software delivery. Establish engineering principles and guardrails that … delivery. Automation of operational processes and platform lifecycle management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. ...