2,676 to 2,700 of 2,830 Remote/Hybrid Observability Jobs

Staff Software Engineer, AI Platform (US - Remote or Calgary)

Hiring Organisation
Syndio
Location
Little Rock, Arkansas, United States
Employment Type
Permanent
Salary
USD Annual
product chat, Slack, Teams) serving our product lines, built on syndi-api: an agent runtime owning orchestration, tool calling, memory, RAG, evals, and observability, live in production today. The platform is young (v1 shipped this quarter), moving fast, and designed around a clear operating model: product teams contribute domain content … eval system - offline gates and online detector/judge evals written onto production traces - and its growth as product teams adopt it. Production operation: observability (Datadog LLM Obs), incident response, the reliability of an agent surface real customers use. The platform's contribution seams: reviewing product teams' tool wrappers ...

Staff Software Engineer, AI Platform (US - Remote or Calgary)

Hiring Organisation
Syndio
Location
Bismarck, North Dakota, United States
Employment Type
Permanent
Salary
USD Annual
product chat, Slack, Teams) serving our product lines, built on syndi-api: an agent runtime owning orchestration, tool calling, memory, RAG, evals, and observability, live in production today. The platform is young (v1 shipped this quarter), moving fast, and designed around a clear operating model: product teams contribute domain content … eval system - offline gates and online detector/judge evals written onto production traces - and its growth as product teams adopt it. Production operation: observability (Datadog LLM Obs), incident response, the reliability of an agent surface real customers use. The platform's contribution seams: reviewing product teams' tool wrappers ...

Staff Software Engineer, AI Platform (US - Remote or Calgary)

Hiring Organisation
Syndio
Location
Cedar Rapids, Iowa, United States
Employment Type
Permanent
Salary
USD Annual
product chat, Slack, Teams) serving our product lines, built on syndi-api: an agent runtime owning orchestration, tool calling, memory, RAG, evals, and observability, live in production today. The platform is young (v1 shipped this quarter), moving fast, and designed around a clear operating model: product teams contribute domain content … eval system - offline gates and online detector/judge evals written onto production traces - and its growth as product teams adopt it. Production operation: observability (Datadog LLM Obs), incident response, the reliability of an agent surface real customers use. The platform's contribution seams: reviewing product teams' tool wrappers ...

Staff Software Engineer, AI Platform (US - Remote or Calgary)

Hiring Organisation
Syndio
Location
El Paso, Texas, United States
Employment Type
Permanent
Salary
USD Annual
product chat, Slack, Teams) serving our product lines, built on syndi-api: an agent runtime owning orchestration, tool calling, memory, RAG, evals, and observability, live in production today. The platform is young (v1 shipped this quarter), moving fast, and designed around a clear operating model: product teams contribute domain content … eval system - offline gates and online detector/judge evals written onto production traces - and its growth as product teams adopt it. Production operation: observability (Datadog LLM Obs), incident response, the reliability of an agent surface real customers use. The platform's contribution seams: reviewing product teams' tool wrappers ...

Staff Software Engineer, AI Platform (US - Remote or Calgary)

Hiring Organisation
Syndio
Location
Las Vegas, Nevada, United States
Employment Type
Permanent
Salary
USD Annual
product chat, Slack, Teams) serving our product lines, built on syndi-api: an agent runtime owning orchestration, tool calling, memory, RAG, evals, and observability, live in production today. The platform is young (v1 shipped this quarter), moving fast, and designed around a clear operating model: product teams contribute domain content … eval system - offline gates and online detector/judge evals written onto production traces - and its growth as product teams adopt it. Production operation: observability (Datadog LLM Obs), incident response, the reliability of an agent surface real customers use. The platform's contribution seams: reviewing product teams' tool wrappers ...

Staff Software Engineer, AI Platform (US - Remote or Calgary)

Hiring Organisation
Syndio
Location
Overland Park, Kansas, United States
Employment Type
Permanent
Salary
USD Annual
product chat, Slack, Teams) serving our product lines, built on syndi-api: an agent runtime owning orchestration, tool calling, memory, RAG, evals, and observability, live in production today. The platform is young (v1 shipped this quarter), moving fast, and designed around a clear operating model: product teams contribute domain content … eval system - offline gates and online detector/judge evals written onto production traces - and its growth as product teams adopt it. Production operation: observability (Datadog LLM Obs), incident response, the reliability of an agent surface real customers use. The platform's contribution seams: reviewing product teams' tool wrappers ...

Staff Software Engineer, AI Platform (US - Remote or Calgary)

Hiring Organisation
Syndio
Location
Santa Clara, California, United States
Employment Type
Permanent
Salary
USD Annual
product chat, Slack, Teams) serving our product lines, built on syndi-api: an agent runtime owning orchestration, tool calling, memory, RAG, evals, and observability, live in production today. The platform is young (v1 shipped this quarter), moving fast, and designed around a clear operating model: product teams contribute domain content … eval system - offline gates and online detector/judge evals written onto production traces - and its growth as product teams adopt it. Production operation: observability (Datadog LLM Obs), incident response, the reliability of an agent surface real customers use. The platform's contribution seams: reviewing product teams' tool wrappers ...

Staff Software Engineer, AI Platform (US - Remote or Calgary)

Hiring Organisation
Syndio
Location
Saint Paul, Minnesota, United States
Employment Type
Permanent
Salary
USD Annual
product chat, Slack, Teams) serving our product lines, built on syndi-api: an agent runtime owning orchestration, tool calling, memory, RAG, evals, and observability, live in production today. The platform is young (v1 shipped this quarter), moving fast, and designed around a clear operating model: product teams contribute domain content … eval system - offline gates and online detector/judge evals written onto production traces - and its growth as product teams adopt it. Production operation: observability (Datadog LLM Obs), incident response, the reliability of an agent surface real customers use. The platform's contribution seams: reviewing product teams' tool wrappers ...

Staff Software Engineer, AI Platform (US - Remote or Calgary)

Hiring Organisation
Syndio
Location
Salt Lake City, Utah, United States
Employment Type
Permanent
Salary
USD Annual
product chat, Slack, Teams) serving our product lines, built on syndi-api: an agent runtime owning orchestration, tool calling, memory, RAG, evals, and observability, live in production today. The platform is young (v1 shipped this quarter), moving fast, and designed around a clear operating model: product teams contribute domain content … eval system - offline gates and online detector/judge evals written onto production traces - and its growth as product teams adopt it. Production operation: observability (Datadog LLM Obs), incident response, the reliability of an agent surface real customers use. The platform's contribution seams: reviewing product teams' tool wrappers ...

Staff Software Engineer, AI Platform (US - Remote or Calgary)

Hiring Organisation
Syndio
Location
Sioux Falls, South Dakota, United States
Employment Type
Permanent
Salary
USD Annual
product chat, Slack, Teams) serving our product lines, built on syndi-api: an agent runtime owning orchestration, tool calling, memory, RAG, evals, and observability, live in production today. The platform is young (v1 shipped this quarter), moving fast, and designed around a clear operating model: product teams contribute domain content … eval system - offline gates and online detector/judge evals written onto production traces - and its growth as product teams adopt it. Production operation: observability (Datadog LLM Obs), incident response, the reliability of an agent surface real customers use. The platform's contribution seams: reviewing product teams' tool wrappers ...

Staff Software Engineer, AI Platform (US - Remote or Calgary)

Hiring Organisation
Syndio
Location
Rapid City, South Dakota, United States
Employment Type
Permanent
Salary
USD Annual
product chat, Slack, Teams) serving our product lines, built on syndi-api: an agent runtime owning orchestration, tool calling, memory, RAG, evals, and observability, live in production today. The platform is young (v1 shipped this quarter), moving fast, and designed around a clear operating model: product teams contribute domain content … eval system - offline gates and online detector/judge evals written onto production traces - and its growth as product teams adopt it. Production operation: observability (Datadog LLM Obs), incident response, the reliability of an agent surface real customers use. The platform's contribution seams: reviewing product teams' tool wrappers ...

Principal Data Engineer

Location
City Of London, England, United Kingdom
follow the identical pattern so they are handover-ready by design. Drive data quality as a first-class, firm-wide concern: establish data contracts, observability, SLA/SLO monitoring, and automated alerting and remediation across ingestion and transformation layers, and hold squads to those standards. Act as the senior technical … with the ability to set standards, conduct code and design reviews, and grow engineers’ capabilities Strong grasp of data quality practices: data contracts, pipeline observability, SLA/SLO definition, and automated alerting and remediation Solid understanding of SQL transformation patterns and modern tooling such as dbt, alongside experience managing ingestion ...

Staff Engineer

Location
Greater London, England, United Kingdom
mobile technologies, build proof‐of‐concepts, and create roadmaps for adoption. Operational Excellence: Champion the "you build it, you run it" philosophy - driving pipeline observability, zero‐downtime deployments, robust monitoring, and alerting for front‐end applications. Resilience & Performance: Design solution‐wide UI resilience (graceful degradation, back‐pressure, retry strategies … understanding of SOLID principles, TDD, DDD, and software design patterns . Experience designing for resilience and performance at scale (graceful degradation, zero‐downtime deployments, observability, performance budgets). Experience with Azure development and cloud‐based infrastructure , including CI/CD and pipeline observability. Experience owning threat modelling and security hardening ...

Data & Analytics Senior Director, Data Solutions & Architecture Data & Analytics London

Location
Greater London, England, United Kingdom
Anthropic Claude, OpenAI and other relevant enterprise platforms. Understand the foundations required for safe and scalable AI delivery, including identity, permissions, environments, data lineage, observability, evaluation, auditability, LLM access patterns and responsible AI controls. Work closely with Data Solutions & Architecture colleagues to ensure AI use cases are built … including structured data, unstructured data, semantic layers, vector search, permissions and governance. Knowledge of safe and scalable AI delivery, including identity, access, environments, observability, evaluation, auditability and responsible AI controls. Experience working with senior clients and translating technical concepts into commercial language. A track record of tying AI and data ...

Staff Reliability Engineer (Full Stack)

Location
Greater London, England, United Kingdom
Native). Lead technical problem-solving during incidents: coordinate response, diagnose root causes, communicate status, and drive to resolution. Build and evolve monitoring/observability (dashboards, alerts, tracing, logging) that enables fast detection and diagnosis. Drive post‐incident reviews (blameless) and ensure learnings become durable fixes (tech changes, runbooks, automation … comfort working across services and APIs. Proven incident response leadership: on-call participation, triage, mitigation, and root‐cause analysis (RCA) with follow‐through. Solid observability skills: practical experience with logging/metrics/tracing and turning signals into actionable alerts and dashboards. Experience collaborating with mobile teams and understanding mobilebackend ...

Staff Software Engineer - Marketplace

Hiring Organisation
Zipline
Location
San Francisco, California, United States
Employment Type
Permanent
Salary
USD Annual
clear SLAs, backpressure handling, and schema/version migration strategies, while serving as the technical lead for critical partner onboarding and incident response. • Maintain observability and service reliability by defining SLOs, building monitoring and dashboards (Grafana/Honeycomb), and owning alerting thresholds, runbooks, and on-call rotations for assigned services. … more of Go, Python, or scalable backend languages; experience with React or similar for merchant-facing apps; familiarity with Kafka, gRPC, PostgreSQL, AWS, observability tools (Grafana, Honeycomb), and build tooling (Bazel) is required for day-one productivity. • Operating intensity and logistics: able to support occasional off-hours launches and urgent ...

Senior Firmware Engineer

Hiring Organisation
Eight Sleep
Location
San Francisco, California, United States
Employment Type
Permanent
Salary
USD Annual
Job Description Job Description Join the Sleep Fitness Movement At Eight Sleep, we're on a mission to fuel human potential through optimal sleep. As the world's first sleep fitness company, we're redefining ...

Senior Infrastructure Software Engineer

Hiring Organisation
Hired Recruiters
Location
Milwaukee, Wisconsin, United States
Employment Type
Permanent
Salary
USD Annual
Job Description Job Description Bicycle Health - Senior Infrastructure Software Engineer SUMMARY Company Stage: Series A 32.3 Employer Tech Stack: React, React Native, TypeScript, GraphQL, GCP Acceptable Tech Background: Distributed Systems, Docker, Kubernetes, Node, TypeScript, Terraform ...

Senior Infrastructure Software Engineer

Hiring Organisation
Hired Recruiters
Location
Honolulu, Hawaii, United States
Employment Type
Permanent
Salary
USD Annual
Job Description Job Description Bicycle Health - Senior Infrastructure Software Engineer SUMMARY Company Stage: Series A 32.3 Employer Tech Stack: React, React Native, TypeScript, GraphQL, GCP Acceptable Tech Background: Distributed Systems, Docker, Kubernetes, Node, TypeScript, Terraform ...

Senior Infrastructure Software Engineer

Hiring Organisation
Hired Recruiters
Location
Cleveland, Ohio, United States
Employment Type
Permanent
Salary
USD Annual
Job Description Job Description Bicycle Health - Senior Infrastructure Software Engineer SUMMARY Company Stage: Series A 32.3 Employer Tech Stack: React, React Native, TypeScript, GraphQL, GCP Acceptable Tech Background: Distributed Systems, Docker, Kubernetes, Node, TypeScript, Terraform ...

Senior Infrastructure Software Engineer

Hiring Organisation
Hired Recruiters
Location
Boston, Massachusetts, United States
Employment Type
Permanent
Salary
USD Annual
Job Description Job Description Bicycle Health - Senior Infrastructure Software Engineer SUMMARY Company Stage: Series A 32.3 Employer Tech Stack: React, React Native, TypeScript, GraphQL, GCP Acceptable Tech Background: Distributed Systems, Docker, Kubernetes, Node, TypeScript, Terraform ...

Senior Infrastructure Software Engineer

Hiring Organisation
Hired Recruiters
Location
Grand Rapids, Michigan, United States
Employment Type
Permanent
Salary
USD Annual
Job Description Job Description Bicycle Health - Senior Infrastructure Software Engineer SUMMARY Company Stage: Series A 32.3 Employer Tech Stack: React, React Native, TypeScript, GraphQL, GCP Acceptable Tech Background: Distributed Systems, Docker, Kubernetes, Node, TypeScript, Terraform ...

Senior Infrastructure Software Engineer

Hiring Organisation
Hired Recruiters
Location
Atlanta, Georgia, United States
Employment Type
Permanent
Salary
USD Annual
Job Description Job Description Bicycle Health - Senior Infrastructure Software Engineer SUMMARY Company Stage: Series A 32.3 Employer Tech Stack: React, React Native, TypeScript, GraphQL, GCP Acceptable Tech Background: Distributed Systems, Docker, Kubernetes, Node, TypeScript, Terraform ...

Senior Infrastructure Software Engineer

Hiring Organisation
Hired Recruiters
Location
Memphis, Tennessee, United States
Employment Type
Permanent
Salary
USD Annual
Job Description Job Description Bicycle Health - Senior Infrastructure Software Engineer SUMMARY Company Stage: Series A 32.3 Employer Tech Stack: React, React Native, TypeScript, GraphQL, GCP Acceptable Tech Background: Distributed Systems, Docker, Kubernetes, Node, TypeScript, Terraform ...

Senior Infrastructure Software Engineer

Hiring Organisation
Hired Recruiters
Location
Salt Lake City, Utah, United States
Employment Type
Permanent
Salary
USD Annual
Job Description Job Description Bicycle Health - Senior Infrastructure Software Engineer SUMMARY Company Stage: Series A 32.3 Employer Tech Stack: React, React Native, TypeScript, GraphQL, GCP Acceptable Tech Background: Distributed Systems, Docker, Kubernetes, Node, TypeScript, Terraform ...