201 to 225 of 698 Remote Observability Jobs

AI Platform Engineer

Hiring Organisation
Addition
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£475.00 per day
from experimentation into production. Implementing RBAC and OIDC-based access controls. Supporting secure API key issuance, rotation and revocation. Applying high standards across testing, observability and technical documentation. Using evaluation frameworks, golden datasets, structured outputs and fallback paths to manage non-deterministic behaviour. Producing architecture decision records, runbooks ...

Senior Product Engineer

Location
Greater London, England, United Kingdom
optimisation RESTful API design and implementation Git - version control and collaborative workflows Testing strategies (unit, integration, end-to-end) Experience with monitoring, logging, and observability tools Knowledge of scalable system architecture and design patterns Familiarity with modern AI tools and a habit of staying up to date with how they ...

Senior / Staff Machine Learning Scientist May 14, 2026

Location
Glasgow, Scotland, United Kingdom
workflow, plus familiarity with cheminformatics tooling (e.g. RDKit, OpenEye) — or willingness to pick these up. MLOps fluency: experiment tracking, data versioning, model serving, and observability of deployed models. A visible track record in the field — peer‐reviewed publications, open‐source contributions, or public projects that demonstrate your judgement on real ...

Senior Data Analyst, People Analytics

Location
Greater London, England, United Kingdom
modelling in a production context Exposure to AI product development, LLM tooling, or data applications Experience with data quality frameworks, semantic layers, or data observability tooling Prior experience working with HR, people, or workforce data Additional Information Bring all of you to work We create the conditions for high performers ...

Technical Architect (AWS and Microservices) - London, UK

Location
Greater London, England, United Kingdom
ensure implementation of permanent fixesSupport incident problem and change management processesEnsure application availability resilience disaster recovery and business continuity requirements are metDrive proactive monitoring observability alerting and operational automation initiativesEnhancement DeliveryPartner with business and product stakeholders to translate requirements into scalable technical solutionsSupport estimation impact assessments and solution planning ...

Staff Engineer - Web Platform

Location
Greater London, England, United Kingdom
evolve the architecture of Fin's web platform. Define the long-term technical strategy for the web team, focusing on scalability,performance, developer productivity, observability, and system reliability. Collaborate closely with marketing, design, analytics, and data science stakeholders toensure the platform supports their goals with accuracy, performance, and agility. Lead ...

Staff enterprise technology engineer

Location
United Kingdom
RISE, on premise ECCs & BWs, S4 Public, SAP Business AI Platform (BTP) and Joule, integration, data and analytics platforms, engineering toolchains, test automation, observability, application lifecycle management and the wider SAP technology ecosystem. This is deliberately broader than a specialist engineering role. The successful candidate will dynamically take more ownership ...

Machine Learning Platform Engineer

Location
Greater London, England, United Kingdom
move into production with confidence Build and orchestrate analytical/processing pipelines across ingestion, training, inference and post-processing with built-in monitoring and observability Design and implement processes to increase efficiency in your workstreams Identify and integrate tooling (including AI tools) to streamline your work Work closely with cross ...

Staff AI Engineer role (Remote Eligible)

Hiring Organisation
Capital One
Location
New York, United States
Employment Type
Permanent
Salary
USD Annual
develop, test, deploy, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, and more Invent ...

Staff AI Engineer role (Remote Eligible)

Hiring Organisation
Capital One
Location
Richmond, Virginia, United States
Employment Type
Permanent
Salary
USD Annual
develop, test, deploy, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, and more Invent ...

Staff AI Engineer role (Remote Eligible)

Hiring Organisation
Capital One
Location
Mc Lean, Virginia, United States
Employment Type
Permanent
Salary
USD Annual
develop, test, deploy, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, and more Invent ...

Staff AI Engineer - Enterprise Analysis Platform (Remote Eligible)

Hiring Organisation
Capital One
Location
New York, United States
Employment Type
Permanent
Salary
USD Annual
develop, test, deploy, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, and more Invent ...

Staff AI Engineer - Enterprise Analysis Platform (Remote Eligible)

Hiring Organisation
Capital One
Location
Mc Lean, Virginia, United States
Employment Type
Permanent
Salary
USD Annual
develop, test, deploy, and support AI software components including foundation model training, large language model inference, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, and more Invent ...

Director, Agentic AI, Data Science Lead, AI Labs

Hiring Organisation
Hackajob Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent
memory, multi-step reasoning - on large, complex datasets, iterating as findings emerge. Partner with engineering to make solutions production-grade and compliant, with observability, guardrails, and evaluation pipelines in place. Manage and mentor AI scientists, guiding technical approach and career growth to build the bench in AI Labs. Communicate technical ...

Director, Agentic AI, Data Science Lead, AI Labs

Hiring Organisation
Hackajob Ltd
Location
Livingston, Scotland, United Kingdom
memory, multi-step reasoning - on large, complex datasets, iterating as findings emerge. Partner with engineering to make solutions production-grade and compliant, with observability, guardrails, and evaluation pipelines in place. Manage and mentor AI scientists, guiding technical approach and career growth to build the bench in AI Labs. Communicate technical ...

Director, Agentic AI, Data Science Lead, AI Labs

Hiring Organisation
Hackajob Ltd
Location
Broughton, Wales, United Kingdom
memory, multi-step reasoning - on large, complex datasets, iterating as findings emerge. Partner with engineering to make solutions production-grade and compliant, with observability, guardrails, and evaluation pipelines in place. Manage and mentor AI scientists, guiding technical approach and career growth to build the bench in AI Labs. Communicate technical ...

Director, Agentic AI, Data Science Lead, AI Labs

Hiring Organisation
Hackajob Ltd
Location
Dunfermline, Scotland, United Kingdom
memory, multi-step reasoning - on large, complex datasets, iterating as findings emerge. Partner with engineering to make solutions production-grade and compliant, with observability, guardrails, and evaluation pipelines in place. Manage and mentor AI scientists, guiding technical approach and career growth to build the bench in AI Labs. Communicate technical ...

Lead Platform Engineer

Location
City of Westminster, England, United Kingdom
testing, documentation and code organisation. Responsible for evolving those standards over time, protecting core principles while adapting to new tools and approaches. Ensure appropriate observability, logging and error-handling patterns are in place across applications. Responsible for ensuring documentation exists where it adds long-term value, and remains accurate. Problem ...

Site Reliability Engineer III

Location
Belfast City District, Northern Ireland, United Kingdom
Service Discovery (Consul, Vault), and Data Distribution (SFTP/JScape)—to Google Cloud Platform. Manage cluster lifecycles, data replication, RBAC, and workload placement. Observability & Monitoring Fabric: Design, scale, and maintain our observability backbone using tools like OpenTelemetry, Splunk, Prometheus, and Grafana. Establish and continuously improve metrics, logs, alerting strategies, SLIs … Strategic communication skills to translate technical requirements for cross-functional teams, coupled with an eagerness to learn independently and collaboratively. Preferred Qualifications/Desirable Observability Stack: Hands-on experience with telemetry tools such as OpenTelemetry, Splunk, Prometheus, and Grafana. Agile Integration: Comfort working within Agile frameworks and collaborative software development ...

Senior Site Reliability Engineer

Location
Knutsford, England, United Kingdom
drive reliability, scalability and performance across critical banking systems. This role combines hands‐on SRE engineering with technical leadership, with a strong focus on observability, automation, continuous improvement and optimisation. Responsibilities Build and maintain reliable, scalable and secure infrastructure platforms and solutions. Apply SRE and software engineering practices to improve … lead complex troubleshooting and root cause analysis. Develop automation using programming and scripting to reduce manual intervention and improve efficiency. Develop and improve observability, monitoring, instrumentation and performance capabilities. Use data and reliability metrics to drive continuous improvement and optimisation. Lead technical discussions, blameless retrospectives and problem‐solving activities. Work ...

Senior Site Reliability Engineer

Hiring Organisation
GCS
Location
Glasgow, City of Glasgow, United Kingdom
Employment Type
Permanent
Salary
£75000 - £95000/annum Bonus
drive reliability, scalability and performance across critical banking systems. This role combines hands-on SRE engineering with technical leadership, with a strong focus on observability, automation, continuous improvement and optimisation. Responsibilities: * Build and maintain reliable, scalable and secure infrastructure platforms and solutions. * Apply SRE and software engineering practices to improve … lead complex troubleshooting and root cause analysis. * Develop automation using programming and scripting to reduce manual intervention and improve efficiency. * Develop and improve observability, monitoring, instrumentation and performance capabilities. * Use data and reliability metrics to drive continuous improvement and optimisation. * Lead technical discussions, blameless retrospectives and problem-solving activities. * Work ...

Remote SRE: Platform Reliability & Observability

Location
United Kingdom
Orexnova is seeking an experienced SRE/Platform Engineer to join a fully remote UK team. You’ll own incident response, blameless post-mortems and drive reliability improvements across services while partnering with the SRE ...

Platform Engineer

Location
Greater London, England, United Kingdom
data and AI workflows. It’s an excellent opportunity for an experienced Platform/DevOps Engineer to work with cloud, Kubernetes, CI/CD, observability, and emerging AI infrastructure while helping establish scalable, secure, and reliable engineering practices. This is an opportunity to join an innovative, progressive, and collaborative team. … agent orchestration AI Evaluation & Quality: Eval harnesses and golden datasets, LLM-as-judge and human-in-the-loop review, regression suites, and red-teaming Observability & Monitoring: Prometheus, Grafana, Datadog, Splunk, Elastic/ELK, OpenTelemetry, including GenAI tracing and token, latency, and cost telemetry Platform Security & Policy-as-Code: HashiCorp Vault ...

Lead Cloud Platform Engineer (Kubernetes) - Remote

Location
United Kingdom
managed platform services, including capacity planning, performance tuning, cost optimisation, patching, and lifecycle management Partner with software engineering teams to support application deployment, troubleshooting, observability, and platform adoption Monitor platform health and respond to incidents, conducting root cause analysis and implementing preventative improvements Identify technical debt and contribute to platform …/CD principles and experience building and maintaining automated delivery pipelines Experience working with container technologies and cloud‐native architectures Knowledge of observability, monitoring, logging, and incident management practices Strong troubleshooting and problem‐solving skills across infrastructure, platform, and application layers Experience supporting software development teams in deploying and operating ...

Senior DevOps Engineer

Location
United Kingdom
image scanning. Implement and monitor infrastructure and application security controls. Support the organisation's ongoing compliance and certification requirements. Reliability & SRE Establish and maintain observability across distributed systems. Develop proactive monitoring, alerting and performance-tuning strategies. Help maintain service-level objectives and platform availability. Investigate and resolve infrastructure and application … advantageous: MLOps or LLMOps experience. Experience with platforms such as SageMaker, Kubeflow or ZenML . Extensive on-premises Kubernetes deployment experience. Prometheus or comparable observability platforms. AWS Karpenter. AWS Compute Optimizer. Experience operating highly distributed systems. Familiarity with ISO 27001, NIST SSDF, OWASP SAMM or similar security frameworks. Understanding ...