SGLang-Diffusion adds day-0 support for Qwen-Image-2.1
SGLang announced day-0 support for Qwen-Image 2.1 in SGLang-Diffusion. Qwen thanked the project and said the stack now serves the model for text-to-image generation, multi-image editing, and transparent RGBA output.
According to SGLang, native precision can run on a single RTX 4090 24GB with CPU offload, with 1024×1024 generation in 18.7s and image editing in 21.7s at 22.7 GiB peak GPU memory during requests. On an RTX PRO 6000 96GB, SGLang reported 8.0s generation and 9.6s editing.
SGLang said one checkpoint covers those image tasks, with native inference including TP/SP, LoRA, and OpenAI-compatible APIs. The reported timings used 40 denoising steps, one image per request, warmed HTTP latency including PNG output, and no quantization.