Did Codex Reset
GitHub

QQ·微信群

#llm-inference

查看全部AI快訊
TGRSS 訂閱共 1 則

DeepLearning.AI 於 The Batch 詳解 DeepSeek 快取縮減與 Flash 基準測試表現

DeepLearning.AI 於最新一期 The Batch 分析了 DeepSeek 的架構,指出其每個 token 的快取縮減至 890 位元組,比 DeepSeek-V1 縮小了 437 倍。分析亦指出,當輸入從 4K 擴展至 1M token 時,每個輸出 token 的運算量增加了 25%,並指出 Flash 於 AA Index 基準測試的得分超越 V4-Pro。