951 to 975 of 1,849 Permanent Observability Jobs

.NET Solution Architect

Hiring Organisation
HTC Global Services Inc
Location
Miami, Florida, United States
Employment Type
Permanent
Salary
USD Annual
tools. Understanding of Responsible AI and AI security principles. Preferred Skills TOGAF or cloud certifications preferred. Experience with AI-assisted development tools. Experience with observability and monitoring tools such as: Datadog Dynatrace Knowledge of AIOps and automated operational frameworks. Experience in ZeroOps/self-healing platform implementations. Education Bachelor … Exposure to GenAI-assisted SDLC workflows. Experience in platform engineering and reusable accelerator frameworks. Knowledge of enterprise integration ecosystems. Experience with AI-driven monitoring, observability, and self-healing platforms. Soft Skills Excellent communication and stakeholder management skills. Strong problem-solving and analytical abilities. Ability to lead distributed teams. Strong presentation ...

Senior Reliability Engineer

Hiring Organisation
Fitch Group
Location
Greater London, United Kingdom
Employment Type
Full Time
someone who is curious about the evolving role of AI in infrastructure engineering, someone who actively explores how AI-assisted tooling, automation, and intelligent observability can raise the bar for reliability and developer experience. You will collaborate closely with global development and engineering teams to deliver reliable, resilient, and high … efficiency Identify, contain, and mitigate risk across all cloud environments, maintaining a robust security posture for infrastructure and applications Implement proactive monitoring and observability practices to detect and prevent issues before they impact users Develop and maintain automation and tooling solutions, including AI-assisted approaches to reduce toil and accelerate ...

Forward Deployed Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
plus.* Hands-on experience with cloud platforms (AWS), Docker and Kubernetes; CI/CD in GitLab (or equivalent); infrastructure-as-code (e.g. Terraform); observability and monitoring stacks.* Solid understanding of database systems, SQL, data modelling, ETL pipelines, REST/gRPC APIs and microservices architecture; identity and access management with OIDC …/SAML and Azure Entra ID.* Experience with LLMs, RAG systems, prompt engineering and AI evaluation frameworks; familiarity with MLOps, model deployment and AI observability/guardrails; working use of AI-assisted development tools (e.g. Claude Code).* Familiarity with one or more of: trading, ERP or treasury platforms; workflow ...

Senior Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
operational efficiency.* Infrastructure-as-Code (IaC): Own and evolve our declarative infrastructure using Terraform for cloud resources and Helm for Kubernetes application deployment.* Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification.* Reliability & Performance: Define, measure, and enforce Service … programming language, preferably Python, for automation and tool development.**Tooling & Concepts*** CI/CD: Experience setting up and maintaining modern CI/CD pipelines.* Observability: Practical experience implementing and managing monitoring and logging tools.* Networking: Solid understanding of TCP/IP, load balancing, DNS, and cloud-native networking within Kubernetes. ...

Java Developer - Security & Intelligence

Hiring Organisation
Jobleads-UK
Location
Gloucester, England, United Kingdom
solutions meet stringent performance, reliability and security requirements. Contribute to architectural design, code reviews, automated testing and continuous improvement activities. Implement logging, monitoring and observability to support the operation of production systems. Collaborate closely with multi‐disciplinary teams including engineers, data specialists and operational stakeholders. Skills Required Strong experience developing … backend systems for data‐intensive or mission‐critical applications. Experience working in DevSecOps environments, including Docker, Kubernetes, CI/CD pipelines, automated testing and observability tooling. Ability to work directly with stakeholders to understand requirements and translate them into robust, secure and scalable software solutions. Strong collaboration and communication skills ...

Java Developer - Security & Intelligence

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
solutions meet stringent performance, reliability and security requirements. Contribute to architectural design, code reviews, automated testing and continuous improvement activities. Implement logging, monitoring and observability to support the operation of production systems. Collaborate closely with multi‐disciplinary teams including engineers, data specialists and operational stakeholders. Skills Required Strong experience developing … backend systems for data‐intensive or mission‐critical applications. Experience working in DevSecOps environments, including Docker, Kubernetes, CI/CD pipelines, automated testing and observability tooling. Ability to work directly with stakeholders to understand requirements and translate them into robust, secure and scalable software solutions. Strong collaboration and communication skills ...

Senior DevOps Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
productivity and experience across the SDLC Cost optimise infra through rightsizing and creating episodic, on-demand environments Proactively monitor for security and reliability using Observability tooling including SIEM, APM, tracing, infrastructure metrics, logs and dashboards Durably engineer away toil You will be a great fit here if you: Are passionate … infrastructure continuously using CI/CD tools such as GitlabCI, CircleCI, Github Actions, and GitOps using ArgoCD, FluxCD Troubleshooting and debugging applications using Observability tooling across microservices and serverless applications such as Splunk, DataDog Managing ephemer secrets and credentials using Hashicorp Vault Managing least privileged access to cloud resources using ...

Sr. Distinguished Machine Learning Engineer (Remote-Eligible)

Hiring Organisation
Capital One
Location
Mc Lean, Virginia, United States
Employment Type
Permanent
Salary
USD Annual
Sr. Distinguished Machine Learning Engineer (Remote-Eligible) Overview: At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine ...

Founding Lead SRE - Remote, Equity

Hiring Organisation
Jobleads-UK
Location
United Kingdom
company. You’ll be a technical and operational leader for production services, mentoring others and setting on-call standards. You’ll own disaster recovery, observability, and incident response while working on our Cloud Application Platform, Kubernetes on AWS, and interconnected engineering teams. #J-18808-Ljbffr ...

Senior Cloud Platform Architect - IaC & FinOps Leader

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
deliver speed, security, cost efficiency, and reliability. You will mentor engineers, drive policy‐based governance, and collaborate with SRE and Security to ensure observability and compliance from day one. #J-18808-Ljbffr ...

Senior Lead SRE & DevOps - Reliability Architect

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
JPMorgan Chase & Co. in Glasgow seeks a Senior Lead Site Reliability/DevOps Engineer to enhance reliability, observability, and performance of critical platforms. You will lead OpenTelemetry pipelines, guide incidents, and drive migrations across hybrid environments to secure scalable systems. Role requires deep cloud-native experience, mastery of monitoring tools ...

(senior) Infrastructure & Devops Engineer (m/w/d)

Hiring Organisation
iVentureGroup GmbH
Location
Hammerbrook, Hamburg, Germany
Employment Type
Permanent
Salary
EUR Annual
Verantwortung für unseren operativen IT-Betrieb (24/7), während du gleichzeitig moderne Plattform-Initiativen vorantreibst. Ob Kubernetes-Cluster, CI/CD-Pipelines oder Observability - du bist in deinem Element . click apply for full job details ...

Senior Lead AI Infrastructure & Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
serving platforms. You will own reliability, performance, and cost efficiency of AI inference stacks, operating across cloud, on-prem, and containerized environments with strong observability and incident response practices. You will lead capacity planning, implement automated remediation, and drive secure, robust software delivery with enterprise-grade tooling and governance. #J ...

Lead Java Developer — Real-Time Risk & Cloud (Hybrid)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
full lifecycle from design to production support, integrating new analytics and data sets across global teams. The role emphasizes scalable microservices, streaming data, and observability with ELK, Prometheus and Grafana. Hybrid work model and competitive benefits are offered. #J-18808-Ljbffr ...

Network Automation Engineer - Low-Latency Finance Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
that enables rapid provisioning and reliable operations. This role focuses on network automation at scale, using Python, Ansible, Terraform, and CI/CD, plus observability, on-call duties, and collaboration with security and platform teams. #J-18808-Ljbffr ...

Senior Platform Engineer - Cloud & Automation

Hiring Organisation
Jobleads-UK
Location
Gaydon, England, United Kingdom
premise systems at Gaydon, Warwickshire. You will drive performance, security and cost improvements while simplifying complex environments. You will lead automation, improve observability and establish standards with cross-functional teams, shaping a cloud-first IT landscape that supports business innovation and resilience. #J-18808-Ljbffr ...

Remote AWS SRE: Build Resilient Cloud Platforms

Hiring Organisation
Jobleads-UK
Location
England, United Kingdom
join a globally operating AI-driven cloud platform team. This fully remote UK role involves maintaining production systems on AWS, implementing automation and observability, and partnering with software, platform, cloud and security engineers to improve reliability. You will handle 24/7 incidents, build resilient cloud services, and drive continuous ...

Senior Platform Engineer: Cloud, Automation & Reliability

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
reliable solutions supporting business-critical systems across the organisation. You’ll lead automation, drive engineering best practices, and collaborate with engineers to improve deployment, observability, and reliability. This role emphasizes mentoring and delivering high-quality, maintainable software in a modern #J-18808-Ljbffr ...

Senior Product Analyst - Market Data & Risk Analytics

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
delivery within Risk Analytics. You will translate business needs into data mappings, user stories, and acceptance criteria, while supporting APIs, data pipelines, and observability across teams. The role requires experience with SQL/Python, data quality, and agile delivery, and involves collaboration with product, engineering, QA, and client-facing #J ...

Platform Engineer: AI Biotech Infra for Scalable Discovery

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
enabled biology lab. This hands-on role covers Kubernetes operations, cloud infrastructure, and production deployment across ML pipelines. You will own CI/CD, observability, and migrations, while collaborating with ML, backend, and product teams to deliver scalable, secure, and robust systems for accelerated drug discovery. #J-18808-Ljbffr ...

Senior LLM Infra & Reliability Lead

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
Lead Software Engineer to design and operate scalable AI infrastructure for reliable LLM serving at production scale. This role emphasizes cloud-based deployments, Kubernetes, observability, and cost‐aware performance tuning. You will own reliability, performance, and cost efficiency of the LLM platform end-to-end, manage diverse hosting environments including ...

Senior AI Platform Engineer – LLM Inference Backend

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
backend services for LLM inference and scalable production systems. You will work on routing, batching, scheduling, streaming responses, and quota management while improving APIs, observability, and reliability across the platform. You will explore model architectures, tokenization costs, and GPU utilization, collaborating with product teams and using enterprise AI tooling ...

Senior Backend Engineer, LLM Inference Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
increases throughput, and maximizes GPU utilization. In this role based in Greater London, you will work in an agile team, learn about model architectures, observability, CI/CD, and secure coding practices, and collaborate across engineering, product, and operations to #J-18808-Ljbffr ...

Senior Real-Time Data Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
deliver reliable, scalable systems. You’ll work on event‐driven data processing with Flink and Kafka, develop Java and .NET APIs, and drive observability and cost‐efficient operations while #J-18808-Ljbffr ...

Platform Engineering Lead — Remote

Hiring Organisation
Jobleads-UK
Location
Exeter, England, United Kingdom
cloud-first transformation from an Azure PaaS .NET estate to a modern AI-first platform. You will own cloud infrastructure, CI/CD, observability, security and the developer platform, guiding a DevOps team into a high‐impact capability. This hands-on leadership role combines technical delivery with people leadership, reporting ...