example, a billion embeddings produced roughly 20 cheaper and visual models served roughly 72% cheaper). We also own core platform surfaces including the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution. About the RoleYou will join a small, high-leverage team building production infrastructure for Generative … open-weights serving stack — real-time GPU endpoints, high-throughput batch inference, and fine-tuning (SFT/DPO/LoRA) — alongside the LLM Gateway, Agent Gateway, evals infrastructure, guardrails, and cost attribution. Design scalable, high-performance systems for model serving, batch inference, GPU autoscaling, and fine-tuning that power real ...