Together AI adds canary rollouts to Dedicated Model Inference
Dedicated Model Inference can now upgrade the model behind a live endpoint without downtime. Traffic shifts in gated steps, and metric gates can pause a rollout if p95 latency or the error rate worsens.