Tencent Hunyuan highlights Hy4 preview compressed to 214 GiB
Tencent Hunyuan (@TencentHunyuan) quoted Zhihu Frontier with the comment "Less size, same intelligence," pointing to an explainer on compressing Hy4 preview.
In the quoted post, Zhihu Frontier (@ZhihuFrontier) said contributor yghstill of Tencent Hunyuan's quantization team described packing the 770B-parameter model by shrinking weights from roughly 1.5TB to 214 GiB. The parameter count remains 770B; the compression changes how those weights are represented. The explainer names the Sherry quantization algorithm, STQ1_0 storage format, and MIX-STQ1_0 mixed-precision allocation, and says the complete mixed-precision model averages about 2.38 bits per weight, with aggressive compression concentrated on expert weights.
The quoted account says the team implemented STQ1_0 CUDA kernels in a patched llama.cpp build and reports nearly unchanged MRCR retrieval performance with a small decline in math. Against UD-IQ1_M at a similar bit budget, it says the mixed-precision model led on the reported evaluations, including a gain of more than five points on MRCR. Tencent Hunyuan's own post does not provide those measurements.