1,476 to 1,500 of 4,209 Observability Jobs

Senior Software Engineer

Hiring Organisation
The Portfolio Group
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£90000/annum
Establish automated testing across backend and frontend applications, including unit, contract and end-to-end testing. Work with the platform engineering team on deployment, observability, logging, tracing and operational readiness. Act as the technical owner for the application and integration layer, making and documenting key architectural decisions. Provide technical guidance … such as Lambda, ECS, API Gateway, S3, CloudFront, Cognito and IAM. Experience designing and operating distributed or event-driven systems. A strong understanding of observability, testing and CI/CD practices. Experience working with data platforms or stores such as MongoDB, OpenSearch or Databricks. Experience integrating internal systems and third ...

Senior Software Engineer

Hiring Organisation
The Portfolio Group
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£90,000 per annum
Establish automated testing across backend and frontend applications, including unit, contract and end-to-end testing. Work with the platform engineering team on deployment, observability, logging, tracing and operational readiness. Act as the technical owner for the application and integration layer, making and documenting key architectural decisions. Provide technical guidance … such as Lambda, ECS, API Gateway, S3, CloudFront, Cognito and IAM. Experience designing and operating distributed or event-driven systems. A strong understanding of observability, testing and CI/CD practices. Experience working with data platforms or stores such as MongoDB, OpenSearch or Databricks. Experience integrating internal systems and third ...

Vice President, Production Services Application Support

Hiring Organisation
Hackajob Ltd
Location
Manchester, UK
Lead technical coordination during major incidents, helping drive rapid diagnosis, recovery, stakeholder communication, and root cause remediation. Drive continuous improvement initiatives focused on automation, observability, service reliability, operational efficiency, and reduction of manual processes. Evaluate production risks associated with application releases, infrastructure changes, and platform enhancements to ensure safe … resolve complex technical issues under pressure. Deep understanding of enterprise application architecture, distributed systems, cloud technologies, middleware, databases, and infrastructure components. Experience with monitoring, observability, automation, and operational tooling used to support highly available production platforms. Strong analytical and problem-solving skills with the ability to identify root causes ...

Senior Backend Engineer - Asset Sales

Location
Greater London, England, United Kingdom
Modern C# stack : Distributed C# and .NET microservices Cloud & orchestration : Hosted on Azure using Kubernetes Architecture : Event-driven, supporting products used at significant scale Observability : Grafana, Azure Application Insights, logs, traces, and metrics AI tooling : Claude and other AI tools used throughout the engineering workflow — design exploration, code generation … want engineers who tinker — experimenting with new tools, agents, and workflows, and sharing what works Guardrails as we accelerate : Automated tests, SLOs, alerting, observability, and deployment safeguards around everything we ship Own it beyond the pull request : Design for idempotency, retries, out-of-order events, and failure modes, and know ...

Embedded Software Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, UK
Employment Type
Full-time
security best practices across the device lifecycleEdge and cloud integration: integrate devices with cloud IoT platforms and backend services; improve telemetry, health monitoring and observability; support reliable field operation and debug fleet issuesCompliance and testing: support testing and validation for safety, EMC, radio and related requirements; prepare test plans …/CD workflows for embedded softwarePractical debugging experience using lab and software tools such as logic analysers, protocol analysers, network sniffers or observability platformsStrong problem-solving skills and the ability to work across hardware and software boundariesBonus: low-power wireless (Thread, Zigbee, BLE); cloud IoT services on AWS, Azure ...

AI Platform Support Engineer (EMEA)

Location
Greater London, England, United Kingdom
combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in. We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle … post incident reviews and operational improvements Build internal tooling, automation, documentation, and runbooks Partner closely with infrastructure, networking, and platform engineering teams Help improve observability, operational visibility, and troubleshooting workflows Improve the customer experience through better processes and technical guidance What This Role Is Not This is not a traditional ...

Embedded Software Engineer

Location
Greater London, England, United Kingdom
best practices across the device lifecycle Edge and cloud integration: integrate devices with cloud IoT platforms and backend services; improve telemetry, health monitoring and observability; support reliable field operation and debug fleet issues Compliance and testing: support testing and validation for safety, EMC, radio and related requirements; prepare test plans …/CD workflows for embedded software Practical debugging experience using lab and software tools such as logic analysers, protocol analysers, network sniffers or observability platforms Strong problem-solving skills and the ability to work across hardware and software boundaries Bonus: low-power wireless (Thread, Zigbee, BLE); cloud IoT services ...

Lead Software Engineer - Cloud

Location
Glasgow, Scotland, United Kingdom
usability, and self-service for engineering consumers Create and maintain delivery workflows that standardize engineering practices and reduce operational toil across the platform Improve observability and operational readiness through monitoring, logging, tracing, alerting, runbook development, and on-call practices Partner with security and risk stakeholders to implement secure-by-default … Familiarity with infrastructure-as-code and automation practices, including tools such as Terraform, Helm, Kustomize, Argo CD, Flux, or continuous integration systems Experience with observability stacks and site reliability engineering practices, including service level indicators and objectives, incident response, and post-incident reviews Exposure to regulated environments and implementing security ...

Senior Full Stack Engineer - Lyst Shop (11 Month FTC - Maternity Cover)

Location
Greater London, England, United Kingdom
rely heavily on experimentation to validate ideas and guide decisions. Technical Excellence: You will help maintain a high bar for code quality, testing, observability, and system reliability. You’ll contribute to architectural discussions, improve developer experience, and proactively address technical debt where needed. Team Contribution: You will mentor and support … working relationships across Product, Design, QA, Analytics, and Engineering teams while actively participating in team ceremonies and technical discussions. Technical Impact: Improve the stability, observability, and maintainability of our systems through better monitoring, resilient code, and thoughtful testing practices. Growth & Ownership: Gain confidence working across our platform and infrastructure while ...

Infrastructure Engineer

Location
Greater London, England, United Kingdom
tenant environments for enterprise customers. Developer Velocity: Architect CI/CD pipelines (GitHub Actions, Docker) that allow our team to ship safely and quickly. Observability & Reliability: Instrument the stack with OpenTelemetry and Datadog to ensure we detect issues before our users do. AI Performance Tuning: Tune autoscaling and network routing … Haves: Experience with ECS, container orchestration, and distributed task queues (Celery/SQS). Strong Python skills. Familiarity with the Datadog/Grafana observability stack. A background in offensive security, CTFs, or AI/ML infrastructure. What We Offer Competitive Salary + Significant Equity: We want you to have true ...

Security Platform Engineer

Hiring Organisation
IBM SIXworks Limited
Location
Farnborough, Hampshire, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
agents Hands-on experience with Kubernetes Experience managing and administering SIEM tooling (e.g. Splunk) Experience deploying vulnerability scanning and analysis tooling (e.g. Nessus) Kubernetes observability (e.g. fluentbit) Familiarity with container security principles and tools Scripting or automation skills (e.g. Python, Bash, or similar) Understanding of security frameworks and best practices … Knowledge of configuring SIEM tooling. Basic understanding of threat frameworks, such as ATT&CK. Experience with additional SIEM or observability platforms Experience with Microsoft Defender Experience with DevSecOps practices and pipeline security Certifications such as CKA, CKAD, CISSP, CEH, or similar Exposure to threat modelling and security architecture design Just ...

Senior AI / Agentic Engineer London

Location
Greater London, England, United Kingdom
Agentic ecosystem, responsible for the high-level design choices that define how agents run at PhysicsX. You will cover topics such as: Agent Observability: Own the implementation to enforce deep tracing, granular cost tracking, and observability across the lifecycle. Agent Deployment: Deliver an intuitive deployment lifecycle which simplifies questions around … behalf of users in a regulated enterprise environment. The Tech Stack Core Platform: Python (Primary), Go or TypeScript (Secondary), Kubernetes, Docker, Terraform. Observability & Evals: OTel, LangSmith, Arize, Braintrust. Who You Are An Architect at Heart: You have strong, reasoned opinions on Durable Execution vs. Standard Async, Vector Search vs. Keyword ...

Security Platform Engineer: Build Secure Infra & CI/CD

Location
Farnborough, England, United Kingdom
agents Hands-on experience with Kubernetes Experience managing and administering SIEM tooling (e.g. Splunk) Experience deploying vulnerability scanning and analysis tooling (e.g. Nessus) Kubernetes observability (e.g. fluentbit) Familiarity with container security principles and tools Scripting or automation skills (e.g. Python, Bash, or similar) Understanding of security frameworks and best practices … Knowledge of configuring SIEM tooling. Basic understanding of threat frameworks, such as ATT&CK. Experience with additional SIEM or observability platforms Experience with Microsoft Defender Experience with DevSecOps practices and pipeline security Certifications such as CKA, CKAD, CISSP, CEH, or similar Exposure to threat modelling and security architecture design Just ...

Staff AI Engineer

Hiring Organisation
Capital One
Location
New York, United States
Employment Type
Permanent
Salary
USD Annual
software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state … vision and the long term roadmap of foundational AI systems at Capital One. Set the technical direction for enterprise-wide AI architecture - unifying tooling, observability, and deployment standards across teams Own the design and integration of model routing, caching, and orchestration systems that support hybrid and multi-model workloads Champion ...

Staff AI Engineer

Hiring Organisation
Capital One
Location
Mc Lean, Virginia, United States
Employment Type
Permanent
Salary
USD Annual
software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state … vision and the long term roadmap of foundational AI systems at Capital One. Set the technical direction for enterprise-wide AI architecture - unifying tooling, observability, and deployment standards across teams Own the design and integration of model routing, caching, and orchestration systems that support hybrid and multi-model workloads Champion ...

Senior Software Engineer - London

Location
Greater London, England, United Kingdom
shorter path each time. Setting the standard for how a benchmark enters the framework, and building the checks that enforce it. Scheduling and observability, so cluster capacity isn't left idle while evaluation jobs queue. Reproducible results across trials, so a release decision rests on numbers that hold. Whatever stack … benchmarks, and started fixing what slows the framework down. By 6 months one part of it is yours, for example scaling the runs, observability, or a group of related benchmarks, and a release will have gone out on your numbers. By 12 months you'll know the design and trade ...

Staff AI Engineer (Remote Eligible)

Hiring Organisation
Capital One
Location
New York, United States
Employment Type
Permanent
Salary
USD Annual
software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state … vision and the long term roadmap of foundational AI systems at Capital One. Set the technical direction for enterprise-wide AI architecture - unifying tooling, observability, and deployment standards across teams Own the design and integration of model routing, caching, and orchestration systems that support hybrid and multi-model workloads Champion ...

Staff AI Engineer role (Remote Eligible)

Hiring Organisation
Capital One
Location
Mc Lean, Virginia, United States
Employment Type
Permanent
Salary
USD Annual
software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state … vision and the long term roadmap of foundational AI systems at Capital One. Set the technical direction for enterprise-wide AI architecture - unifying tooling, observability, and deployment standards across teams Own the design and integration of model routing, caching, and orchestration systems that support hybrid and multi-model workloads Champion ...

Senior AI Engineer - Agentic AI

Location
Greater London, England, United Kingdom
that allow complex AI workflows to operate securely and efficiently at scale. You will be responsible for developing advanced orchestration capabilities, implementing evaluation and observability tooling and embedding enterprise controls for compliance and safety. If you are passionate about innovating with AI in real-world applications and scaling intelligent systems … ensuring graceful degradation and retries Apply enterprise security and governance practices including RBAC, prompt safety checks, traceability and secrets management Implement evaluation pipelines and observability frameworks using tools such as Langfuse, Arize or OpenTelemetry Contribute to architectural design decisions, code reviews and engineering standards for platform development Requirements Bachelor ...

Staff Python Engineer (ML)

Location
City Of London, England, United Kingdom
apps in a service architecture. Furthering Developer Experience (DevEx) by mentoring others in writing code that is intuitive, clear, and easy to test Developing observability for new and existing ML applications and GenAI/LLM integrations , making use of the Grafana Stack (Prometheus, Loki, Tempo) Develop integrations and services that … Backend-Engineering Experience owning projects from start to finish, including speccing, architecture, development, testing, deployment, release and monitoring Strong skills in building maintainable tests, observability and tracing systems. Knowledge of best practices for performance optimisation, memory management. Familiarity with Kubernetes , Docker and other cloud infrastructure, ops and containerised tools. Strong ...

Junior DevOps Engineer

Location
Greater London, England, United Kingdom
/CD processes, automate manual processes to create a self-service environment for our developers, and maintain platform uptime SLAs by improving our observability stack. We use infrastructure-as-code to maintain our platform on AWS, so familiarity with common AWS Services (RDS, S3, ECS, EC2 etc.) and Terraform …/CD pipelines through Jenkins/Github Actions and other technology Support the Development and AI Engineering Teams - help troubleshoot their issues Bring observability through dashboards, alerting and log aggregation Run incident analysis and post mortems Environment management: ephe...dev/staging/prod What you’ll bring A strong foundation ...

Senior Network Engineer, Studios

Location
Greater London, England, United Kingdom
modelled in a NetBox source of truth, configuration is deployed from code through pipelines rather than hand‐edited, and streaming telemetry feeds the observability stack. Changes are expected to land in days, not months, in a facility where the network carries live content around the clock. The estate spans dedicated … defect. Config as code: build and run the automation that deploys configuration from intent through CI pipelines (Python, Ansible, Git); eliminate hand‐edits. Observability: develop monitoring built on streaming telemetry (gNMI/gRPC), flow analysis and modern tooling (Grafana, Zabbix class), serving the 24/7 operations teams as your ...

Software Engineer / AI Engineer

Location
Greater London, England, United Kingdom
within a defined problem, building and testing tool use, retrieval pipelines and agent workflows, integrating AI capabilities into enterprise systems, and contributing to evaluation, observability and guardrails. You will hold a high bar on code quality, flag risks and blockers early, and work alongside host‐function stakeholders to make sure … agentic AI solutions to production standard within a defined technical approach. Implement and test tool use, retrieval pipelines, and agent workflows. Contribute to evaluation, observability and guardrails for agentic systems. Integrate AI capabilities into existing enterprise workflows and systems. Maintain high code quality and documentation so patterns can be reused. ...

Staff AI Engineer - Enterprise Analysis Platform (Remote Eligible)

Hiring Organisation
Capital One
Location
New York, United States
Employment Type
Permanent
Salary
USD Annual
software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state … vision and the long term roadmap of foundational AI systems at Capital One. Set the technical direction for enterprise-wide AI architecture - unifying tooling, observability, and deployment standards across teams Own the design and integration of model routing, caching, and orchestration systems that support hybrid and multi-model workloads Champion ...

Staff AI Engineer - Enterprise Analysis Platform (Remote Eligible)

Hiring Organisation
Capital One
Location
Mc Lean, Virginia, United States
Employment Type
Permanent
Salary
USD Annual
software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state … vision and the long term roadmap of foundational AI systems at Capital One. Set the technical direction for enterprise-wide AI architecture - unifying tooling, observability, and deployment standards across teams Own the design and integration of model routing, caching, and orchestration systems that support hybrid and multi-model workloads Champion ...