vLLM integrates PyNvVideoCodec to offload video decoding to GPU NVDEC
vLLM said it has integrated PyNvVideoCodec to offload video decoding from the CPU to the GPU's NVDEC.
According to the project, the result is 2x+ throughput at 8×H100, with the CPU bottleneck gone. It said this matters for video captioning at scale (AV training, metadata) and ships with CUDA vLLM releases.