specific needs. Dive deep into the entire stack, from Kubernetes and container orchestration, through gRPC‐based service communication, to the performance tuning of ONNX‐based inference on GPU‐accelerated hardware. Write clean, efficient, and rigorously tested code. We value simplicity, correctness, and peer review. What you'll bring … challenges of managing the lifecycle of models in a multi‐tenant, high‐availability system. Familiarity with building ML inference services, model serialization (e.g., ONNX), and GPU programming (CUDA). You've built or worked on custom storage or job‐queueing systems before and have the scars to prove it. Maybe ...