2,501 to 2,525 of 4,558 Observability Jobs

AI Engineer 4 (MLX, Agentic AI, Gen AI platform Services)

Hiring Organisation
Capital One
Location
Mc Lean, Virginia, United States
Employment Type
Permanent
Salary
USD Annual
software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state … long term roadmap of foundational AI systems at Capital One. Own the end-to-end architecture for complex AI systems - ensuring maintainability, observability, and ethical alignment Define and maintain service-level objectives (SLOs) for AI reliability, including latency, uptime, and model performance drift Collaborate with infrastructure engineering to optimize ...

AI Engineer 4 (AI Foundations)

Hiring Organisation
Capital One
Location
New York, United States
Employment Type
Permanent
Salary
USD Annual
software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state … long term roadmap of foundational AI systems at Capital One. Own the end-to-end architecture for complex AI systems - ensuring maintainability, observability, and ethical alignment Define and maintain service-level objectives (SLOs) for AI reliability, including latency, uptime, and model performance drift Collaborate with infrastructure engineering to optimize ...

AI Engineer 4 (AI Foundations: LLM Customization, Finetuning, Reinforcement Learning)

Hiring Organisation
Capital One
Location
Mc Lean, Virginia, United States
Employment Type
Permanent
Salary
USD Annual
software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state … long term roadmap of foundational AI systems at Capital One. Own the end-to-end architecture for complex AI systems - ensuring maintainability, observability, and ethical alignment Define and maintain service-level objectives (SLOs) for AI reliability, including latency, uptime, and model performance drift Collaborate with infrastructure engineering to optimize ...

AI Engineer 4 (AI Foundations)

Hiring Organisation
Capital One
Location
Mc Lean, Virginia, United States
Employment Type
Permanent
Salary
USD Annual
software components including foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability, etc. Leverage a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state … long term roadmap of foundational AI systems at Capital One. Own the end-to-end architecture for complex AI systems - ensuring maintainability, observability, and ethical alignment Define and maintain service-level objectives (SLOs) for AI reliability, including latency, uptime, and model performance drift Collaborate with infrastructure engineering to optimize ...

Bioinformatics Scientist/Engineer – DRAGEN Array

Location
United Kingdom
bioinformatics solutions, integrating established community tools alongside novel methods. You will also contribute to software engineering best practices, including testing automation, CI/CD, observability, and software lifecycle management, helping to ensure DRAGEN Array software is reliable, scalable, and ready for production deployment across local and cloud-based environments.**What …/CD workflows to support continuous integration, validation, and reliable delivery of DRAGEN Array software.* Implement and evolve telemetry, logging, and monitoring to improve observability, diagnostics, and operational robustness.* Execute and improve Software Lifecycle (SLC) practices, including requirements traceability, design documentation, verification, and validation.* Use tools such as JAMA (requirements ...

Senior Machine Learning Engineer Software engineering London

Location
Greater London, England, United Kingdom
proficiency in writing clear, production‐ready Python code. Experience with production ML models (online or offline) and standard MLOps practices. Experience with monitoring and observability of production systems, with a strong sense of ownership. Experience with training and operating models on Databricks. Familiarity with cloud‐based application development (AWS & Azure ...

Cloud Security Control Engineer (AWS)

Hiring Organisation
Diana Duggan UK Limited
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£445.00 per day
/Rego), DevOps/CI/CD and security controls experience. Strong commercial experience in AWS cloud engineering. Experience working with monitoring, logging or observability platforms such as Splunk. Experience with Infrastructure as Code, particularly Terraform. Strong hands-on Python development and scripting experience. Proven experience designing, implementing and supporting ...

C# Software Engineer Full Stack

Hiring Organisation
Client Server
Location
Bracknell, Berkshire, UK
Employment Type
Full-time
architecture and technical decisions within your domain, collaborate closely with Product, Design and Compliance to take features from idea to production and improve reliability, observability, testing and production performance, modernising legacy services and contributing to wider architectural direction. Location/WFH:You can work from home most of the time ...

Lead AI Engineer Python LLM - Tech Consultancy

Hiring Organisation
Client Server
Location
London, UK
Employment Type
Full-time
STEM disciplines You have hands-on experience deploying LLMs and multi-modal models at scale in productionYou have a strong understanding of scalable MLOps, observability and cloud-native AI deploymentYou have strong Python coding skills and API development skillsYou're collaborative and pragmatic with advanced stakeholder communication, problem-solving ...

Senior Full Stack Engineer C

Hiring Organisation
Client Server
Location
Bracknell, Berkshire, UK
Employment Type
Full-time
architecture and technical decisions within your domain, collaborate closely with Product, Design and Compliance to take features from idea to production and improve reliability, observability, testing and production performance, modernising legacy services and contributing to wider architectural direction. Location/WFH:You can work from home most of the time ...

C# Developer .Net API Full Stack

Hiring Organisation
Client Server
Location
Bracknell, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
architecture and technical decisions within your domain, collaborate closely with Product, Design and Compliance to take features from idea to production and improve reliability, observability, testing and production performance, modernising legacy services and contributing to wider architectural direction. Location/WFH: You can work from home most of the time ...

Ping Directory Lead (AD-to-Ping Directory Migration) (Hybrid 2/3 Days)

Hiring Organisation
Bitsoft International, Inc
Location
Malvern, Pennsylvania, United States
Employment Type
Any
Salary
USD Annual
APIs/SCIM/IAM/PAM Linux/Java JVM Tuning/Python/PowerShell/Shell Scripting Splunk/SIEM/Monitoring & Observability Backup & Restore/Capacity Planning/Performance Tuning Security, Compliance & Least Privilege Production Incident Management ...

Backend Software Engineer

Location
Greater London, England, United Kingdom
retrieval, and evals — not one-off prompt demos Improve document extraction and auto-itemization: quality, tax fields, latency, cost, and provider failover Instrument production: observability, quality dashboards, chat satisfaction, extraction accuracy vs. human correction Partner with Product and Design on employee, admin, and mobile AI surfaces (including unified travel … expenses, payments, ERP, tax, or similarly high-stakes data. Clear written design, strong code review, and bias to production quality (tests, CI/CD, observability, incident ownership). Based in Berlin or London (or willing to relocate). Collaboration across US, Israel, and India time zones. Nice to have: document ...

Backend Software Engineer Python LLM - Finance

Hiring Organisation
Client Server
Location
East London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
development You have hands-on experience deploying LLMs and multi-modal models at scale in production You have a strong understanding of scalable MLOps, observability and cloud-native AI deployment You're collaborative and pragmatic with advanced stakeholder communication, problem-solving and project management skills in Agile environments Experience with ...

Java Developer (React)

Location
Greater London, England, United Kingdom
/CD tools (Jenkins, TeamCity) and deployment tools (Ansible) Knowledge of front office trading systems (FX, Rates, Credit, or derivatives) Experience with monitoring and observability tools Strong problem‐solving mindset with a proactive approach #J-18808-Ljbffr ...

Platform Engineer

Location
Slough, England, United Kingdom
Core Focus) Build secure, scalable Azure and GCP cloud platforms. Drive IaC adoption using Terraform and cloud-native tooling. Implement automation, CI/CD, observability, and reliability engineering practices. Define standards for networking, identity, security, monitoring, and resilience. Improve developer experience through platform consistency, automation, and self-service. Governance, Security ...

Platform Engineer

Location
Greater London, England, United Kingdom
Core Focus) Build secure, scalable Azure and GCP cloud platforms. Drive IaC adoption using Terraform and cloud-native tooling. Implement automation, CI/CD, observability, and reliability engineering practices. Define standards for networking, identity, security, monitoring, and resilience. Improve developer experience through platform consistency, automation, and self-service. Governance, Security ...

Principal FDE - Software Engineer

Hiring Organisation
Microsoft
Location
United Kingdom, UK
Employment Type
Full-time
working closely with engineers, applied scientists, TPMs, designers, and operational teams. Apply strong software engineering practices including code reviews, automated testing, CI/CD, observability, documentation, and operational readiness. Build reusable components, accelerators, and engineering patterns that enable successful solutions to scale across teams and transformation scenarios. Learn and adapt … abreast of current developments. Proactively seeks new knowledge and adapts to new trends, technical solutions, and patterns that will improve the availability, reliability, efficiency, observability, and performance of products while also driving consistency in monitoring and operations at scale and shares knowledge with other engineers. Applies, extrapolates, and helps ...

Dataiku Solution Architect

Hiring Organisation
Everforth Quinnox
Location
City of London, London, United Kingdom
Employment Type
Permanent
dashboards use controlled, reconciled, and traceable data from the governed platform. Define access, refresh, performance, lineage, and reconciliation standards for reporting solutions. Support curve observability and the monitoring of data quality, source availability, processing status, and workflow completion. Nonfunctional Architecture Define nonfunctional requirements for performance, scalability, security, availability, resiliency, recoverability … observability, maintainability, and supportability. Design monitoring and alerting across Dataiku, Power Automate, PostgreSQL or Amazon RDS, integrations, WebApps, and reporting components. Establish recovery patterns for failed source deliveries, workflow errors, data-quality issues, integration failures, and interrupted processing. Define capacity and performance considerations for regional processing, historical replay, concurrent users ...

Managing Engineer – Database, Platform

Location
Belfast, Northern Ireland, United Kingdom
product engineering, SRE, security, and cloud teams to define SLAs/SLOs, incident response, root cause analysis, and risk mitigation. Establish monitoring, alerting, and observability for database health and performance. Drive proactive incident prevention. Define and enforce standards, best practices, and governance for database access, security, backups, and compliance. Participate ...

Principal AI Security Software Engineer

Location
Greater London, England, United Kingdom
software factory concepts. Experience delivering on cloud platforms, ideally AWS, with containers and Infrastructure as Code. A solid grounding in secure by default engineering, observability and production operability, with practical knowledge of AI security practices such as the OWASP Top 10 for LLM and Generative AI applications and OWASP guidance ...

Software Engineer

Location
Cheltenham, England, United Kingdom
work from technical documentation and requirements. Practical experience using AI-assisted development tools. Microservices. AWS Lambda, ECS or EC2. CI/CD. Docker. Observability, logging and monitoring. Secure coding and resilient application development. We value diversity and equal opportunity. All applicants will receive consideration for employment without regard to race ...

Site Reliability Engineer

Hiring Organisation
Morgan McKinley
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
Salary negotiable
point for complex Windows Server OS internals, Active Directory forests, Group Policies, Kerberos authentication, and core domain services. Site Reliability Engineering (SRE): Implement proactive observability, monitoring, automated recovery workflows, and performance optimization to enforce continuous service availability and operational resilience. Architecture & Modernization: Partner with engineering, cloud, and security teams ...

ML Engineer

Hiring Organisation
VIQU Limited
Location
London, UK
Employment Type
Full-time
into batch and real-time environments through APIs, scheduled workflows and production pipelinesManage model versioning, promotion and rollback throughout the ML lifecycleImplement monitoring and observability across production models, including model and data drift, performance alerts and loggingDevelop automated retraining processes to maintain model performance and reliabilityWork closely with Data Engineering ...

AI Assisted Java Developer

Hiring Organisation
Square One Resources
Location
Leeds, West Yorkshire, United Kingdom
Employment Type
Contract
Contract Rate
£550 - £600/day
assisted development tools and spec-driven engineering practices to improve productivity, code quality, and delivery effectiveness. Champion engineering excellence through automated testing, code reviews, observability, maintainability, and continuous improvement. Collaborate closely with product managers, designers, QA engineers, and fellow developers to solve complex customer problems. Support and mentor other engineers ...